Information processing device, information processing system, information processing method and control program
The information processing system improves user convenience by using learned models to evaluate and arrange document information based on user-specified extraction criteria, enhancing the efficiency and accuracy of document retrieval.
Patent Information
- Application Number
- JP2023223706
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing information processing systems that output specific document information from multiple documents lack convenience for users in efficiently extracting relevant content.
An information processing apparatus and system that utilizes learned models to evaluate document information, allowing users to specify extraction numbers or ratios, and arranges output in descending order of evaluation values, with optional correction based on keywords and attributes, and includes storage for field-specific extraction settings.
Enhances user convenience by efficiently extracting high-quality document information relevant to the user's needs, reducing labor and improving the accuracy and efficiency of document search.
Smart Images

Figure 2025105267000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing system, an information processing method, and a control program.
Background Art
[0002] Conventionally, various documents such as papers have been created by various researchers, developers, etc. When a researcher or developer wants to obtain specific knowledge, they can efficiently obtain the knowledge by referring to the documents created by other researchers or developers. In recent years, an information processing system has been developed that outputs specific document information from among a plurality of document information regarding each document so that a user can search for a desired document from among a plurality of documents.
[0003] Non-Patent Document 1 describes the development and evaluation of RobotReviewer, a machine learning system that automatically evaluates bias in clinical trials. This system determines the risk of bias from a PDF-formatted trial report and extracts sentences that support these determinations.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In an information processing system that outputs specific document information from among a plurality of document information, it is required to further improve the convenience for the user.
[0006] The purpose of the information processing apparatus, information processing system, information processing method, and control program is to enable improvement of the convenience for the user.
Means for Solving the Problem
[0007] The information processing apparatus according to the embodiment includes an acquisition unit that acquires a plurality of document information, a specific unit that specifies the evaluation value of each document information by inputting the plurality of document information to a learned model that has been pre-learned to output the evaluation value of a predetermined document information when the predetermined document information is input, a reception unit that receives a designation of the number of extractions or extraction ratio by the user, an extraction unit that extracts the document information of the number of extractions or extraction ratio in descending order of the evaluation value from among the plurality of document information, and an output unit that outputs data in which the extracted document information is arranged in the order of the evaluation value.
[0008] The information processing apparatus according to the embodiment includes an acquisition unit that acquires a plurality of document information, a specific unit that specifies the evaluation value of the first part of each document information by inputting the plurality of document information to a first learned model that has been pre-learned to output the evaluation value of the first part of a predetermined document information when the predetermined document information is input, a reception unit that receives a designation of the number of extractions or extraction ratio by the user, an extraction unit that extracts the document information of the number of extractions or extraction ratio in descending order of the evaluation value from among the plurality of document information, a second specific unit that specifies the second evaluation value of the second part of each of the extracted document information by inputting the extracted document information to a second learned model that has been pre-learned to output the second evaluation value of the second part of a predetermined document information when the predetermined document information is input, and an output unit that outputs data in which the extracted document information is arranged in the order of the second evaluation value.
[0009] In the information processing apparatus according to the embodiment, it further has a storage unit in which the extraction number or extraction ratio is stored for each of a plurality of fields. The acquisition unit further acquires the field to which a plurality of document information belongs, and when the reception unit has not received the designation of the extraction number or extraction ratio by the user, it is preferable that the extraction unit extracts the document information of the extraction number or extraction ratio corresponding to the field to which the plurality of document information belongs from among the plurality of document information.
[0010] In the information processing apparatus according to the embodiment, when the reception unit has not received the designation of the extraction number or extraction ratio by the user and the extraction number or extraction ratio corresponding to the field to which the plurality of document information belongs is not stored in the storage unit, it is preferable that the extraction unit calculates the statistical value of the extraction number or extraction ratio of each field stored in the storage unit and extracts the document information corresponding to the statistical value from among the plurality of document information.
[0011] In the information processing apparatus according to the embodiment, it is preferable that it further has a setting unit that receives the designation of the document information in the data by the user and sets the extraction number or extraction ratio corresponding to the field to which the plurality of document information belongs based on the designated document information.
[0012] In the information processing apparatus according to the embodiment, it is preferable that the acquisition unit further acquires a keyword, and the specifying unit corrects the evaluation value of each document information according to whether the keyword is included in each of the plurality of document information.
[0013] In the information processing apparatus according to the embodiment, it is preferable that the acquisition unit further acquires the attributes of each of the plurality of document information, and the specifying unit corrects the evaluation value of each document information based on the attributes of each of the plurality of document information.
[0014] The information processing system according to the embodiment is an information processing system having a first information processing device and a second information processing device. The first information processing device includes an acquisition unit that acquires a plurality of document information, and a specifying unit that specifies the evaluation value of each document information by inputting the plurality of document information into a learned model that has been pre-learned to output the evaluation value of the predetermined document information when the predetermined document information is input. The second information processing device includes a reception unit that receives a designation of the number of extractions or the extraction ratio by the user, an extraction unit that extracts the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from the plurality of document information, and an output unit that outputs data in which the extracted document information is arranged in the order of the evaluation value.
[0015] The information processing system according to the embodiment is an information processing system having a first information processing device and a second information processing device. The first information processing device includes an acquisition unit that acquires a plurality of document information, and a specifying unit that specifies the evaluation value of the first part of each document information by inputting the plurality of document information into a first learned model that has been pre-learned to output the evaluation value of the first part of the predetermined document information when the predetermined document information is input. The second information processing device includes a reception unit that receives a designation of the number of extractions or the extraction ratio by the user, an extraction unit that extracts the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from the plurality of document information, a second specifying unit that specifies the second evaluation value of the second part of each extracted document information by inputting the extracted document information into a second learned model that has been pre-learned to output the second evaluation value of the second part of the predetermined document information when the predetermined document information is input, and an output unit that outputs data in which the extracted document information is arranged in the order of the second evaluation value.
[0016] The information processing method according to the embodiment acquires a plurality of document information, specifies the evaluation value of each document information by inputting each of the plurality of document information into a learned model that has been pre-learned to output the evaluation value of the predetermined document information when the predetermined document information is input, receives a designation of the number of extractions or the extraction ratio by the user, extracts the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from the plurality of document information, and outputs data in which the extracted document information is arranged in the order of the evaluation value from an output unit.
[0017] The information processing method according to the embodiment acquires a plurality of document information, and inputs the plurality of document information into a first pre-trained model that outputs an evaluation value of a first part of a predetermined document information when the predetermined document information is input, to identify the evaluation value of the first part of each document information, receives a designation of the number of extractions or the extraction ratio by the user, extracts the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from among the plurality of document information, inputs the extracted document information into a second pre-trained model that outputs a second evaluation value of a second part of the predetermined document information when the predetermined document information is input, to identify the second evaluation value of the second part of each of the extracted document information, and outputs from an output unit data in which the extracted document information is arranged in the order of the second evaluation value.
[0018] The control program according to the embodiment is a control program for an information processing apparatus, which causes the information processing apparatus to acquire a plurality of document information, input each of the plurality of document information into a pre-trained model that outputs an evaluation value of a predetermined document information when the predetermined document information is input, to identify the evaluation value of each document information, receive a designation of the number of extractions or the extraction ratio by the user, extract the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from among the plurality of document information, and output from an output unit data in which the extracted document information is arranged in the order of the evaluation value.
[0019] The control program according to the embodiment is a control program for an information processing apparatus, which acquires a plurality of document information, inputs the plurality of document information to a first pre-trained model pre-trained to output an evaluation value of a first part of a predetermined document information when the predetermined document information is input, identifies the evaluation value of the first part of each document information, receives a designation of the number of extractions or the extraction ratio by the user, extracts the document information of the number of extractions or the extraction ratio in descending order of the evaluation value from among the plurality of document information, inputs the extracted document information to a second pre-trained model pre-trained to output a second evaluation value of a second part of the predetermined document information when the predetermined document information is input, identifies the second evaluation value of the second part of each of the extracted document information, and causes the information processing apparatus to output, from an output unit, data in which the extracted document information is arranged in the order of the second evaluation value.
Effect of the Invention
[0020] The information processing apparatus, the information processing system, the information processing method, and the control program can improve the convenience for the user.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0022] Hereinafter, an information processing apparatus, an information processing system, an information processing method, and a control program according to one aspect of the embodiment will be described with reference to the drawings. However, note that the technical scope of the present invention is not limited to those embodiments, and extends to the invention described in the claims and equivalents thereof.
[0023] FIG. 1 is a diagram showing a schematic configuration of an information processing system 1 according to an embodiment.
[0024] As shown in FIG. 1, the information processing system 1 includes one or more terminal devices 100 and one or more information processing devices 200 and the like. Each terminal device 100 and each information processing device 200 are communicably connected to each other via a network N. The network N is a wired network such as the Internet or an intranet. The network N may be a wireless network such as a wireless LAN (Local Area Network).
[0025] The information processing system 1 manages document information regarding documents (literature) such as papers, articles, and e-books. The documents are technical documents in various technical fields such as electricity, machinery, chemistry, and biochemistry. For example, the document is a paper regarding the function of foods / food components. The document information is information indicating the attributes and contents of each document. The information processing system 1 is used, for example, in a systematic review. A systematic review is a method of thoroughly investigating documents, removing data biases such as publication bias as much as possible from high-quality research data such as randomized controlled trials, and performing analysis. A systematic review is a method of first setting a question to be clarified, comprehensively collecting studies being conducted regarding that question, scrutinizing the information, and then leading to a conclusion by summarizing the entire information.
[0026] Systematic review includes a first stage of comprehensively collecting information, a second stage of extracting necessary information from comprehensive and extensive information, and a third stage of integrating information to draw a single conclusion. In the first stage, the information processing system 1 comprehensively searches for relevant papers using literature databases such as PubMed and Medical Journal Web. In the second stage, the information processing system 1 performs screening (narrowing down) to appropriately extract papers that conduct research on the questions set from the many collected papers. For example, the information processing system 1 first performs a primary screening to simply determine the necessity of papers using only the information of the title and abstract. Then, the information processing system 1 performs a secondary screening to extract only the necessary papers using the entire paper for the papers determined to be necessary in the primary screening. In the third stage, the information processing system 1 finally summarizes the scientific quality or research results of the individually extracted papers and draws a final conclusion.
[0027] Figure 2 is a diagram showing the schematic configuration of the terminal device 100.
[0028] The terminal device 100 is a device such as a personal computer, a notebook PC, a tablet PC, or a multifunctional mobile phone (so-called smartphone). The terminal device 100 includes a first input device 101, a first display device 102, a first communication device 103, a first storage device 110, a first processing device 120, and the like. The first input device 101, the first display device 102, the first communication device 103, the first storage device 110, and the first processing device 120 are interconnected via a CPU (Central Processing Unit) bus or the like.
[0029] The first input device 101 has a touch panel type input device or an input device such as a keyboard and a mouse, and an interface circuit that acquires signals from the input device, and outputs an operation signal corresponding to the input operation of the user.
[0030] The first display device 102 includes a display such as a liquid crystal display or an organic EL (Electro-Luminescence) display, and an interface circuit that outputs image data to the display, and displays the image data on the display.
[0031] The first communication device 103 has a wired communication interface circuit that complies with a communication protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol). The first communication device 103 communicates and connects with the network N according to a communication standard such as Ethernet (registered trademark). The first communication device 103 sends the data received from the information processing device 200 via the network N to the first processing device 120. Also, the first communication device 103 transmits the data received from the first processing device 120 to the information processing device 200 via the network N. Note that the first communication device 103 may have an antenna for transmitting and receiving radio signals and a wireless communication interface circuit that complies with a communication protocol such as a wireless LAN, and communicate and connect with the network N according to a communication standard such as a wireless LAN.
[0032] The first storage device 110 includes a memory device such as a RAM (Random Access Memory) or a ROM (Read Only Memory), a fixed disk device such as a hard disk, or a portable storage device such as a flexible disk or an optical disk. Also, various computer programs, databases, tables, etc. used for various processes of the terminal device 100 are stored in the first storage device 110. The computer program may be installed in the first storage device 110 using a known setup program or the like from a computer-readable portable recording medium. The portable recording medium is, for example, a CD-ROM (compact disc read only memory), a DVD-ROM (digital versatile disc read only memory), etc. The computer program may be stored in a recording medium possessed by a predetermined server and installed via the network N.
[0033] The first processing device 120 operates based on a program stored in the first storage device 110 in advance. The first processing device 120 is, for example, a CPU. As the first processing device 120, a DSP (digital signal processor), an LSI (large scale integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), etc. may be used. The first processing device 120 reads a computer program stored in the first storage device 110 and operates according to the read computer program. The first processing device 120 is connected to the first input device 101, the first display device 102, the first communication device 103, the first storage device 110, etc., and controls each device. The first processing device 120 transmits a plurality of document information to the terminal device 100 via the first communication device 103. The first processing device 120 receives, via the first communication device 103, the document information extracted by the terminal device 100 and arranged in a predetermined order from the terminal device 100 and displays it on the first display device 102.
[0034] Figure 3 is a diagram showing a schematic configuration of the information processing device 200.
[0035] The information processing device 200 is a device such as a server, a personal computer, or a notebook PC. The information processing device 200 includes a second input device 201, a second display device 202, a second communication device 203, a second storage device 210, a second processing device 220, etc. The second input device 201, the second display device 202, the second communication device 203, the second storage device 210, and the second processing device 220 are interconnected via a CPU bus or the like.
[0036] The second input device 201 includes an input device such as a keyboard and a mouse, and an interface circuit that acquires signals from the input device, and outputs an operation signal according to a user's input operation.
[0037] The second display device 202 is an example of an output unit. The second display device 202 includes a display such as a liquid crystal or an organic EL, and an interface circuit that outputs image data to the display, and displays the image data on the display.
[0038] The second communication device 203 is an example of an output unit. The second communication device 203 has a wired communication interface circuit that conforms to a communication protocol such as TCP / IP. The second communication device 203 communicates and connects with the network N according to a communication standard such as Ethernet (registered trademark). The second communication device 203 sends the data received from the terminal device 100, other information processing devices 200, etc. via the network N to the second processing device 220. The second communication device 203 transmits the data received from the second processing device 220 to the terminal device 100, other information processing devices 200, etc. via the network N. Note that the second communication device 203 may have an antenna that transmits and receives radio signals and a wireless communication interface circuit that conforms to a communication protocol such as a wireless LAN, and communicate and connect with the network N according to a communication standard such as a wireless LAN.
[0039] The second storage device 210 is an example of a storage unit. The second storage device 210 includes a memory device such as a RAM or a ROM, a fixed disk device such as a hard disk, or a portable storage device such as a flexible disk or an optical disk. Also, various computer programs, databases, tables, etc. used for various processes of the information processing device 200 are stored in the second storage device 210. The computer program may be installed in the second storage device 210 using a known setup program or the like from a computer-readable portable recording medium such as a CD-ROM or a DVD-ROM. The computer program may be stored in a recording medium of a predetermined server and installed via the network N.
[0040] The second storage device 210 stores, as data, a document table 211, a condition table 212, an extraction table 213, a first learned model 214, a second learned model 215, and the like. The document table 211 stores document information and the like of each document for a plurality of documents. The condition table 212 stores conditions for document information whose extraction is prohibited from among a plurality of pieces of document information, conditions for correcting evaluation values, and the like. The extraction table 213 stores the number of documents to be extracted or the extraction ratio of document information to be extracted from among a plurality of pieces of document information for each of a plurality of fields. Details of the document table 211, the condition table 212, and the extraction table 213 will be described later. The first learned model 214 and the second learned model 215 are models for specifying evaluation values of document information. The first learned model 214 and the second learned model 215 are generated by the information processing device 200 or another server device.
[0041] The second processing device 220 operates based on a program stored in the second storage device 210 in advance. The second processing device 220 is, for example, a CPU. As the second processing device 220, a DSP, an LSI, an ASIC, an FPGA, or the like may be used. The second processing device 220 is connected to the second communication device 203, the second storage device 210, and the like, and controls each device. The second processing device 220 extracts a predetermined number of pieces of document information from among a plurality of pieces of document information received from the terminal device 100 via the second communication device 203, generates output data in which the extracted document information is arranged in a predetermined order, and transmits the output data to the terminal device 100 via the second communication device 203.
[0042] The second processing device 220 reads a computer program stored in the second storage device 210 and operates according to the read computer program. As a result, the second processing device 220 functions as a reception unit 221, an acquisition unit 222, a preprocessing unit 223, a specification unit 224, a second specification unit 225, an extraction unit 226, an output control unit 227, and a setting unit 228.
[0043] FIG. 4(A) is a schematic diagram showing an example of the data structure of the document table 211.
[0044] In the document table 211, for each of a plurality of documents, identification information (document ID) of each document, document information, etc. are stored. The document information includes attribute information and content information.
[0045] The attribute information indicates the attributes of each document (document information), and includes database information, author information, publication information, type information, language information, etc. The database information indicates the database in which each document is stored and the address of the storage area within that database. The author information indicates the author of each document. The publication information indicates the magazine, book, page, etc. in which each document is published. The type information indicates the type of each document (paper, review, proceedings, etc.). The language information indicates the language (Japanese, English, etc.) in which each document is written.
[0046] The content information indicates the content of each document, and includes title information, abstract information, body text information, memo information, etc. The title information indicates the title of each document. The abstract information indicates the abstract of each document. The body text information indicates the body text of each document (the part other than the title, abstract, and memo). The memo information indicates the memo (references, explanations, etc.) described in the footnotes, etc. of each document.
[0047] FIG. 4(B) is a schematic diagram showing an example of the data structure of the condition table 212.
[0048] In the condition table 212, conditions (exclusion conditions) under which the information processing apparatus 200 is prohibited from extracting document information from among a plurality of pieces of document information, conditions (correction conditions) under which the evaluation value is corrected, etc. are stored. As each condition, it is defined that specific information included in the document information has a specific value. The exclusion condition and / or the correction condition is, for example, that the database information indicates a database other than a specific database, the publication information indicates a book other than a specific book, the language information indicates a language other than a specific language (English, Japanese, etc.), the type information indicates a specific type (proceedings, etc.), the memo information includes specific information, etc. In the example shown in FIG. 4(B), as each condition, a combination of the database information having a specific value and specific information other than the database information included in the document information having a specific value is defined.
[0049] Figure 4(C) is a schematic diagram showing an example of the data structure of the extraction table 213.
[0050] In the extraction table 213, extraction information is stored for each of a plurality of fields. The fields include, for example, the technical fields (electricity, machinery, chemistry, biochemistry, etc.) to which each document belongs and / or the types of each document (papers, reviews, conference proceedings, etc.). The extraction information indicates the number or ratio of document information to be extracted by the information processing apparatus 200 from among a plurality of document information.
[0051] Figure 5 is a flowchart showing an example of the operation of the setting process of the information processing apparatus 200.
[0052] Hereinafter, an example of the operation of the setting process of the information processing apparatus 200 will be described with reference to the flowchart shown in Figure 5. Note that the flowchart of the operation described below is mainly executed by the second processing apparatus 220 in cooperation with each element of the information processing apparatus 200 based on a program stored in the second storage apparatus 210 in advance.
[0053] First, the reception unit 221 waits until it receives a designation by the user of the number or ratio of document information to be extracted by the information processing apparatus 200 from among a plurality of document information (step S101). The reception unit 221 acquires the extraction information indicating the number or ratio of extraction designated by the user using the second input device 201 or the terminal device 100 by receiving the information from the second input device 201 or the second communication device 203.
[0054] When the reception unit 221 receives a designation of the number or ratio of extraction by the user, it sets the extraction information indicating the designated number or ratio of extraction by storing it in the second storage apparatus 210 (step S102), and returns the process to step S101.
[0055] Figure 6 is a flowchart showing an example of the operation of the extraction process of the information processing apparatus 200.
[0056] Next, an example of the operation of the extraction process of the information processing apparatus 200 will be described with reference to the flowchart shown in FIG. 6. Note that the flow of the operation described below is mainly executed by the second processing apparatus 220 in cooperation with each element of the information processing apparatus 200 based on a program stored in advance in the second storage apparatus 210. The extraction process shown in FIG. 6 is executed in parallel with the setting process shown in FIG. 5.
[0057] First, the acquisition unit 222 acquires a plurality of document information, fields, and / or keywords (step S201). The field is the field to which the acquired plurality of document information belongs. The keyword is a term included in the document information that the information processing apparatus 200 should preferentially extract from the plurality of document information. The acquisition unit 222 acquires the plurality of document information, fields, and / or keywords specified by the user using the second input device 201 or the terminal device 100 from the second input device 201 or the second communication device 203. The acquisition unit 222 may acquire not the document information directly specified by the user, but the database information and / or the posted information specified by the user. In that case, the acquisition unit 222 accesses the database indicated by the database information via the second communication device 203 to acquire the corresponding document, and acquires the document information by extracting the portion indicated by the posted information from the acquired document.
[0058] Next, the preprocessing unit 223 executes preprocessing on the plurality of document information acquired by the acquisition unit 222 (step S202). For example, as preprocessing, the preprocessing unit 223 excludes the document information that satisfies the exclusion condition from the candidates to be extracted by the information processing apparatus 200. The preprocessing unit 223 refers to the condition table 212, identifies the document information that satisfies the exclusion condition among the plurality of document information acquired by the acquisition unit 222, and excludes it from the candidates to be extracted by the information processing apparatus 200. Thereby, the information processing apparatus 200 can reduce the number of document information to be analyzed, and can reduce the processing load and processing time of the extraction process.
[0059] Note that the preprocessing unit 223 may refer to the database information, author information, and / or publication information of each document information, identify document information whose content is described in other document information (duplicates other document information), and exclude it from the candidates extracted by the information processing apparatus 200. Thereby, the information processing apparatus 200 can reduce the number of document information to be analyzed, and can reduce the processing load and processing time of the extraction process.
[0060] Further, as preprocessing, the preprocessing unit 223 edits (processes) each document information so that each document information satisfies the input format for the first learned model 214 and / or the second learned model 215. For example, the preprocessing unit 223 extracts specific content information (title information, abstract information, text information, and / or memo information) from each document information as information to be input to the first learned model 214 or the second learned model 215.
[0061] Next, the specifying unit 224 specifies an evaluation value of each document information by inputting each document information acquired by the acquisition unit 222 and preprocessed by the preprocessing unit 223 to the first learned model 214 (step S203). The first learned model 214 is pre-learned to output an evaluation value of the predetermined document information when the predetermined document information is input. The first learned model 214 is learned using learning document information including correct document information that should be determined as necessary document information and incorrect document information that should be determined as unnecessary document information. The correct document information and the incorrect document information are classified by an expert having sufficient knowledge in the field to which each document information belongs by determining whether the quality of each document information is high or low.
[0062] The first learned model 214 is trained to output a higher evaluation value as the degree of similarity between the input document information and any correct document information is higher. The first learned model 214 may be trained to output a lower evaluation value as the degree of similarity between the input document information and each incorrect document information is higher. The degree of similarity is, for example, the degree of match of words, context, etc. included in each document, or the normalized cross-correlation value, inner product, etc. of vectors indicating the distributed representations of each character group included in each document. The first learned model 214 is trained by supervised learning such as XGBoost, Nystroem Kernel SVM Classifier, BERT, etc.
[0063] As the first learned model 214, separate models may be used for each mutually different language (Japanese, English, etc.). In that case, the specifying unit 224 specifies the language in which each document information is described from the language information included in each document information, and inputs each document information to the model corresponding to the specified language. For example, the model corresponding to English is trained with XGBoost, and the model corresponding to Japanese is trained with Nystroem Kernel SVM Classifier. Thereby, the information processing apparatus 200 can calculate the evaluation value of each document information more accurately, and can extract high-quality document information more accurately. Further, the specifying unit 224 may use a known translation technique to translate each document information into the language corresponding to the first learned model 214 and then input it to the first learned model 214.
[0064] Next, the specifying unit 224 corrects the specified evaluation value (step S204). For example, the specifying unit 224 corrects the evaluation value of each document information according to whether the keyword acquired by the acquisition unit 222 is included in each document information. In that case, when the keyword is included in each document information, the specifying unit 224 increases the evaluation value of that document information, and when the keyword is not included in each document information, the specifying unit 224 corrects the evaluation value of that document information so as to decrease the evaluation value. The specifying unit 224 may correct the evaluation value of each document information so that the higher the number of keywords included in each document information, the higher the evaluation value. Thereby, the information processing apparatus 200 can increase the evaluation value of the document information suitable for the user's use, and can more accurately extract the document information suitable for the user's use.
[0065] Further, the specifying unit 224 may correct the evaluation value of each document information based on the attribute of each document information. In that case, the specifying unit 224 specifies the attribute indicated by the attribute information included in the document information acquired by the acquisition unit 222 as the attribute of each document information. The specifying unit 224 refers to the condition table 212, specifies the document information in which the specified attribute satisfies the correction condition among each document information, and corrects the evaluation value of each document information so as to decrease the evaluation value of the specified document information. The specifying unit 224 may correct the evaluation value of each document information so that the higher the number of correction conditions satisfied by each document information, the lower the evaluation value. Thereby, the information processing apparatus 200 can more accurately calculate the evaluation value of each document information based on the attribute of each document, and can more accurately extract appropriate document information.
[0066] Next, the extraction unit 226 determines whether the reception unit 221 has received the user's designation of the extraction number or extraction ratio in the setting process (step S205).
[0067] When the reception unit 221 has received the user's designation of the extraction number or extraction ratio, the extraction unit 226 specifies the extraction information indicating the designated extraction number or extraction ratio (the extraction information set in step S102 of FIG. 5) (step S206), and shifts the process to step S210.
[0068] On the other hand, when the reception unit 221 has not received the specification of the extraction number or extraction ratio by the user, the extraction unit 226 identifies the field acquired by the acquisition unit 222 in step S201, that is, the field to which a plurality of document information belongs. The extraction unit 226 determines whether extraction information corresponding to the identified field is stored (set) in the extraction table 213 (step S207).
[0069] When the extraction information corresponding to the identified field is stored, the extraction unit 226 identifies the extraction information corresponding to the identified field (step S208), and transfers the process to step S210.
[0070] On the other hand, when the extraction information corresponding to the identified field is not stored, the extraction unit 226 calculates the statistical value of the extraction number or extraction ratio indicated by the extraction information of each field stored in the extraction table 213 (step S209), and transfers the process to step S210. The extraction unit 226 calculates the average value, median value, minimum value, or maximum value of the extraction number or extraction ratio indicated by the extraction information of all fields stored in the extraction table 213 as the statistical value. The extraction unit 226 may calculate the weighted average value weighted by the number of learning document information for each field or the number of document information processed by the information processing apparatus 200 so far as the statistical value.
[0071] Next, the extraction unit 226 extracts, from among a plurality of document information, the document information of the extraction number or extraction ratio indicated by the extraction information identified in step S206 or S208, or the document information corresponding to the statistical value calculated in step S210, in descending order of the evaluation value (step S210).
[0072] In this way, when the reception unit 221 has received the specification of the extraction number or extraction ratio by the user, the extraction unit 226 extracts the document information of the extraction number or extraction ratio specified by the user from among a plurality of document information in descending order of the evaluation value. Thereby, the information processing apparatus 200 can extract an appropriate number of document information according to the purpose or use of the user.
[0073] Further, when the reception unit 221 has not received the specification of the extraction number or extraction ratio by the user, the extraction unit 226 extracts document information with an extraction number or extraction ratio corresponding to the field to which the plurality of document information belongs from among the plurality of document information. Thereby, the information processing apparatus 200 can extract an appropriate number of document information according to the field to which the document information belongs.
[0074] Also, when the reception unit 221 has not received the specification of the extraction number or extraction ratio by the user and the extraction number or extraction ratio corresponding to the field to which the plurality of document information belongs is not stored in the second storage device 210, the extraction unit 226 calculates the statistical value of the extraction number or extraction ratio of each field stored in the second storage device 210. Then, the extraction unit 226 extracts document information corresponding to the calculated statistical value from among the plurality of document information. Thereby, even when the appropriate number according to the field to which the document information belongs is unknown, the information processing apparatus 200 can extract an appropriate number of document information.
[0075] Next, the output control unit 227 generates output data in which the document information extracted by the extraction unit 226 is arranged in the order of the evaluation values (step S211). The output control unit 227 generates output data in which each document information is arranged in descending order of the evaluation value. The output control unit 227 may generate output data in which each document information is arranged in ascending order of the evaluation value. The output control unit 227 generates output data so that each document information can be specified by the user using a radio button, a check box, or the like.
[0076] Next, the output control unit 227 outputs the generated output data by transmitting it to the terminal device 100 via the second communication device 203 (step S212). For example, the output control unit 227 transmits the output data to the terminal device 100 that transmitted the plurality of document information in step S201. The first processing device 120 of the terminal device 100 receives the output data from the information processing device 200 via the first communication device 103, and notifies the user by displaying the received output data on the first display device 102. When a plurality of document information is input using the second input device 201 in step S201, the output control unit 227 may output the generated output data by displaying it on the second display device 202.
[0077] Next, the setting unit 228 receives a designation of document information in the output data by the user. Then, the setting unit 228 sets the number of extractions or the extraction ratio corresponding to the field to which the plurality of document information belongs based on the designated document information (step S213), and ends a series of steps. The user browses the document information in the order displayed on the second display device 202 or the first display device 102. The user designates the document information by pressing a radio button or a check box corresponding to the document information in which the necessary information is described using the second input device 201 or the terminal device 100. The setting unit 228 acquires the document information designated by the user using the second input device 201 or the terminal device 100 by receiving it from the second input device 201 or the second communication device 203.
[0078] The setting unit 228 identifies the rank of the document information specified by the user within the output data. The setting unit 228 calculates a numerical value obtained by adding a margin to the identified rank, or a ratio obtained by dividing the numerical value by the total number of document information acquired by the acquisition unit 222 in step S101. The setting unit 228 updates the extraction number or extraction ratio corresponding to the field acquired by the acquisition unit 222 in step S101 in the extraction table 213 to the calculated number or ratio. Thereby, the information processing apparatus 200 can update the number or ratio of document information to be extracted after the next time based on the document information actually selected by the user. Thereafter, the information processing apparatus 200 can generate output data with a reduced number of document information while including the document information required by the user, and can improve the convenience for the user.
[0079] Note that the processes of steps S202, S204, S205, S207, S208, S209, and / or S213 may be omitted.
[0080] Also, each process included in the setting process of FIG. 5 and the extraction process of FIG. 6 may be executed in cooperation by a plurality of information processing apparatuses 200. For example, the processes of steps S201 to S204 in FIG. 6 may be executed by the first information processing apparatus, and the processes of steps S101 to S102 in FIG. 5 and the processes of steps S205 to S213 in FIG. 6 may be executed by a second information processing apparatus different from the first information processing apparatus.
[0081] FIG. 7 is a schematic diagram showing a specific result in which document information in which necessary information is described is manually specified from among a plurality of document information for each of a plurality of fields.
[0082] The worst positive example position in FIG. 7 indicates the ratio corresponding to the position from the head of the group of the document information with the lowest evaluation value among the document information in which the necessary information is described within the group in which a plurality of document information are arranged in descending order of evaluation value. The smaller the worst positive example position, the higher the evaluation value of all the document information in which the necessary information is described, and the larger the worst positive example position, the lower the evaluation value of any of the document information in which the necessary information is described.
[0083] As shown in FIG. 7, the average value of the worst positive example positions in all fields was 15.30%. Therefore, the information processing apparatus 200 can extract most of the document information in which necessary information is described while excluding a large number of document information in which necessary information is not described, by setting the extraction ratio of the document information to 15.30% (or setting the extraction number corresponding thereto). Note that 99% of the document information in which necessary information is described belonged to the range within 34% from the top of the group in which all document information was arranged in descending order of the evaluation value. Therefore, the information processing apparatus 200 can extract 99% of the document information in which necessary information is described while excluding a large number of document information in which necessary information is not described, by setting the extraction ratio of the document information to 34% (or setting the extraction number corresponding thereto). As a result, the user can efficiently select necessary document information from the output data from which a large number of document information in which necessary information is not described has been excluded, and can significantly reduce the labor required for document search. Therefore, the information processing apparatus 200 can improve the convenience for the user.
[0084] Also, as shown in FIG. 7, the worst positive example positions are greatly different for each of a plurality of fields. Therefore, the information processing apparatus 200 can extract a large number of document information in which necessary information is described while excluding a large number of document information in which necessary information is not described, by setting the extraction number or extraction ratio of the document information for each of the plurality of fields. Therefore, the information processing apparatus 200 can further improve the convenience for the user.
[0085] As described in detail above, the information processing apparatus 200 extracts document information having the extraction number or extraction ratio specified by the user in descending order of the evaluation value from among a plurality of document information, and outputs output data in which the extracted document information is arranged in order of the evaluation value. As a result, the user can efficiently select necessary document information from the output data from which unnecessary document information has been excluded, and can significantly reduce the labor required for document search. Therefore, the information processing apparatus 200 has become capable of improving the convenience for the user.
[0086] In addition, in order for the information processing apparatus 200 to extract document information similar to high-quality document information previously selected by experts, the user can select high-quality document information from among a plurality of pieces of document information, and can efficiently obtain the necessary knowledge. Therefore, the information processing apparatus 200 can improve the convenience for the user.
[0087] The division value (recall rate by person) obtained by dividing the number A of papers determined to be necessary for both a specific person and the information processing apparatus 200 and actually necessary by the sum of that number A and the number B of papers determined to be necessary for the specific person but actually unnecessary was 0.756. On the other hand, the division value (recall rate by the information processing apparatus 200) obtained by dividing the number A by the sum of that number A and the number C of papers determined to be necessary for the information processing apparatus 200 but actually unnecessary was 0.554. The recall rate by the information processing apparatus 200 was not inferior to the recall rate by person, indicating that the information processing apparatus 200 has performance not inferior to that of humans.
[0088] FIG. 8 is a flowchart showing an example of the operation of the extraction process of the information processing apparatus 200 according to another embodiment.
[0089] Hereinafter, an example of the operation of the extraction process according to the present embodiment will be described while referring to the flowchart shown in FIG. 8. The flow of the operation described below is mainly executed by the second processing apparatus 220 in cooperation with each element of the information processing apparatus 200 based on a program stored in the second storage apparatus 210 in advance. The extraction process shown in FIG. 8 is executed in parallel with the setting process shown in FIG. 5. The processes in steps S301 to S302, S305 to S310, and S314 to S315 in FIG. 8 are the same as the processes in steps S201 to S202, S205 to S210, and S212 to S213 in FIG. 6, so the description thereof will be omitted. Hereinafter, only the processes in steps S303 to S304 and S311 to S313 will be described.
[0090] In step S303, the specifying unit 224 specifies the evaluation value of the first part of each document information by inputting each document information acquired by the acquisition unit 222 and pre-processed by the pre-processing unit 223 into the first learned model 214 (step S303). The first learned model 214 according to the present embodiment is learned in the same manner as the first learned model 214 used in FIG. 6. However, the first learned model 214 according to the present embodiment is pre-learned to output the evaluation value of the first part of the predetermined document information when the predetermined document information is input. The first part is a part showing the outline of the document information, such as title information, abstract information, or memo information, for example.
[0091] The first learned model 214 is learned to output a higher evaluation value as the degree of similarity between the first part of the input document information and the first part of any correct document information is higher. The first learned model 214 may be learned to output a lower evaluation value as the degree of similarity between the first part of the input document information and the first part of each incorrect document information is higher.
[0092] Next, the specifying unit 224 corrects the specified evaluation value (step S304). The specifying unit 224 corrects the evaluation value in the same manner as the process of step S204. Note that the specifying unit 224 may correct the evaluation value of each document information according to whether the keyword acquired by the acquisition unit 222 is included in the first part of each document information.
[0093] In step S311, the second specifying unit 225 specifies the second evaluation value of the second part of each extracted document information by inputting each document information extracted in step S310 into the second learned model 215 (step S311). The second learned model 215 is learned in the same manner as the first learned model 214. However, the second learned model 215 is pre-learned to output the second evaluation value of the second part of the predetermined document information when the predetermined document information is input. The second part is a part showing the whole of the document information, such as body text information, for example. The second part may be all of the document information. The second part is larger than the first part and includes more detailed information than the first part.
[0094] The second learned model 215 is learned to output a second evaluation value that is higher as the degree of similarity between the input document information and any correct document information is higher. The second learned model 215 may be learned to output a lower second evaluation value as the degree of similarity between the input document information and each incorrect document information is higher.
[0095] Next, the second specifying unit 225 corrects the specified second evaluation value (step S312). The second specifying unit 225 corrects the second evaluation value in the same manner as the process of step S204. Note that the specifying unit 224 may correct the second evaluation value of each document information according to whether the keyword acquired by the acquisition unit 222 is included in the second part of each document information.
[0096] Next, the output control unit 227 generates output data in which the document information extracted by the extraction unit 226 is arranged in the order of the second evaluation values (step S313). The output control unit 227 generates output data in which each document information is arranged in descending order of the second evaluation value. The output control unit 227 may generate output data in which each document information is arranged in ascending order of the second evaluation value. The output control unit 227 generates output data so that each document information can be specified by the user using a radio button, a check box, or the like.
[0097] Note that the processes of steps S302, S304, S305, S307, S308, S309, S312, and / or S315 may be omitted.
[0098] In addition, each process included in the setting process of FIG. 5 and the extraction process of FIG. 8 may be executed in cooperation by a plurality of information processing apparatuses 200. For example, the processes of steps S301 to S304 in FIG. 8 may be executed by a first information processing apparatus, and the processes of steps S101 to S102 in FIG. 5 and the processes of steps S305 to S315 in FIG. 8 may be executed by a second information processing apparatus different from the first information processing apparatus.
[0099] As described in detail above, the information processing apparatus 200 extracts document information of the number of extracts or extraction ratio specified by the user in descending order of the evaluation value from among a plurality of document information, and outputs output data in which the extracted document information is arranged in descending order of the second evaluation value. Thereby, the user can efficiently select necessary document information from the output data from which unnecessary document information has been excluded, and can significantly reduce the labor required for document search. Therefore, the information processing apparatus 200 has become capable of improving the convenience for the user.
[0100] In particular, the information processing apparatus 200 uses only the first part with a small range for all document information, excludes unnecessary document information with a low load and in a short time, and for the remaining small number of document information, uses the second part with a large range to sort each document information in descending order of quality. Thereby, the information processing apparatus 200 has become capable of reducing the processing load and processing time of the extraction process while maintaining the accuracy of the output data.
[0101] Note that the embodiment is not limited to the above. For example, instead of using the learned model stored in its own device, the information processing apparatus 200 may specify the evaluation value and / or the second evaluation value using the learned model stored in another server device or the like. In that case, in step S203 of FIG. 5 and steps S303 and S311 of FIG. 8, the specifying unit 224 or the second specifying unit 225 transmits a plurality of document information to the server device via the second communication device 203. The server device receives a plurality of document information from the information processing apparatus 200, inputs it to the learned model, and transmits the evaluation value or the second evaluation value output from the learned model to the information processing apparatus 200. The specifying unit 224 or the second specifying unit 225 acquires the evaluation value or the second evaluation value by receiving it from the server device via the second communication device 203.
[0102] The information processing apparatus 200 can specify an evaluation value and / or a second evaluation value using the latest learned model updated by the server apparatus by using the learned model stored in another server apparatus. Also, the information processing apparatus 200 can reduce the storage capacity. On the other hand, the information processing apparatus 200 can specify the evaluation value and / or the second evaluation value even in a state where the communication connection with the server apparatus is disconnected by using the learned model stored in its own apparatus. Also, the information processing system 1 can reduce the communication volume between the information processing apparatus 200 and the server apparatus.
Explanation of Signs
[0103] 1 Information processing system, 200 Information processing apparatus, 202 Second display apparatus, 203 Second communication apparatus, 210 Second storage apparatus, 221 Reception unit, 222 Acquisition unit, 223 Preprocessing unit, 224 Specification unit, 225 Second specification unit, 226 Extraction unit, 227 Output control unit, 228 Setting unit
Claims
1. An acquisition unit that acquires a plurality of document information; A specifying unit that specifies an evaluation value of each document information by inputting the plurality of document information into a learned model that has been pre-learned to output an evaluation value of the predetermined document information when the predetermined document information is input; A reception unit that receives a designation of the number of extractions or the extraction ratio by the user; An extraction unit that extracts document information of the number of extractions or the extraction ratio from among the plurality of document information in descending order of the evaluation value; An output unit that outputs data in which the extracted document information is arranged in the order of the evaluation value; An information processing apparatus, characterized by comprising the above.
2. An acquisition unit that acquires a plurality of document information; A specifying unit that specifies an evaluation value of the first part of each document information by inputting the plurality of document information into a first learned model that has been pre-learned to output an evaluation value of the first part of the predetermined document information when the predetermined document information is input; A reception unit that receives a designation of the number of extractions or the extraction ratio by the user; An extraction unit that extracts document information of the number of extractions or the extraction ratio from among the plurality of document information in descending order of the evaluation value; A second specifying unit that specifies a second evaluation value of the second part of each of the extracted document information by inputting the extracted document information into a second learned model that has been pre-learned to output a second evaluation value of the second part of the predetermined document information when the predetermined document information is input; An output unit that outputs data in which the extracted document information is arranged in the order of the second evaluation value; An information processing apparatus, characterized by comprising the above.
3. Further comprising a storage unit in which the number of extractions or the extraction ratio is stored for each of a plurality of fields, The acquisition unit further acquires the field to which the plurality of document information belongs, The extraction unit, when the reception unit has not received a designation of the number of extractions or the extraction ratio by the user, extracts document information of the number of extractions or the extraction ratio corresponding to the field to which the plurality of document information belongs from among the plurality of document information. The information processing apparatus according to claim 1 or 2.
4. When the extraction unit has not received a designation of the number of extractions or the extraction ratio by the user and the number of extractions or the extraction ratio corresponding to the field to which the plurality of document information belongs is not stored in the storage unit, the extraction unit calculates a statistical value of the number of extractions or the extraction ratio of each field stored in the storage unit, and extracts document information corresponding to the statistical value from among the plurality of document information. The information processing apparatus according to claim 3.
5. The information processing apparatus according to claim 3, further comprising a setting unit that receives a designation of document information in the data by a user and sets an extraction number or an extraction ratio corresponding to a field to which the plurality of document information belongs based on the designated document information.
6. The acquisition unit further acquires a keyword, The information processing apparatus according to claim 1 or 2, wherein the specifying unit corrects an evaluation value of each document information based on whether or not the keyword is included in each of the plurality of document information.
7. The acquisition unit further acquires an attribute of each of the plurality of document information, The information processing apparatus according to claim 1 or 2, wherein the specifying unit corrects an evaluation value of each document information based on the attribute of each of the plurality of document information.
8. An information processing system including a first information processing apparatus and a second information processing apparatus, The first information processing apparatus, An acquisition unit that acquires a plurality of document information, A specifying unit that specifies an evaluation value of each document information by inputting the plurality of document information into a learned model that has been pre-learned to output an evaluation value of the predetermined document information when the predetermined document information is input, The second information processing apparatus, A reception unit that receives a designation of an extraction number or an extraction ratio by a user, An extraction unit that extracts document information of the extraction number or the extraction ratio from the plurality of document information in descending order of the evaluation value, An output unit that outputs data in which the extracted document information is arranged in the order of the evaluation value, An information processing system characterized by comprising:
9. An information processing system including a first information processing apparatus and a second information processing apparatus, The first information processing apparatus, An acquisition unit that acquires a plurality of document information, A specifying unit that specifies an evaluation value of a first part of each document information by inputting the plurality of document information into a first learned model that has been pre-learned to output an evaluation value of a first part of the predetermined document information when the predetermined document information is input, The second information processing apparatus, A reception unit that receives a designation of an extraction number or an extraction ratio by a user, An extraction unit that extracts document information of the extraction number or the extraction ratio from the plurality of document information in descending order of the evaluation value, Inputting the extracted document information into a second pre-trained model that has been pre-trained to output a second evaluation value for the second part of the predetermined document information when the predetermined document information is input, to identify the second evaluation value for the second part of each of the extracted document information; An output unit that outputs data in which the extracted document information is arranged in the order of the second evaluation values; An information processing system, characterized by comprising the above.
10. Acquiring a plurality of document information, Inputting each of the plurality of document information into a pre-trained model that has been pre-trained to output an evaluation value of the predetermined document information when the predetermined document information is input, to identify the evaluation value of each document information, Receiving a specification of the number of extractions or extraction ratio by the user, Extracting the document information of the number of extractions or the extraction ratio from among the plurality of document information in descending order of the evaluation value, Outputting, from an output unit, data in which the extracted document information is arranged in the order of the evaluation values. An information processing method, characterized by the above.
11. Acquiring a plurality of document information, Inputting the plurality of document information into a first pre-trained model that has been pre-trained to output an evaluation value of the first part of the predetermined document information when the predetermined document information is input, to identify the evaluation value of the first part of each document information, Receiving a specification of the number of extractions or extraction ratio by the user, Extracting the document information of the number of extractions or the extraction ratio from among the plurality of document information in descending order of the evaluation value, Inputting the extracted document information into a second pre-trained model that has been pre-trained to output a second evaluation value of the second part of the predetermined document information when the predetermined document information is input, to identify the second evaluation value of the second part of each of the extracted document information, Outputting, from an output unit, data in which the extracted document information is arranged in the order of the second evaluation values. An information processing method, characterized by the above.
12. A control program for an information processing apparatus, Acquiring a plurality of document information, Inputting each of the plurality of document information into a pre-trained model that has been pre-trained to output an evaluation value of the predetermined document information when the predetermined document information is input, to identify the evaluation value of each document information, Receiving a specification of the number of extractions or extraction ratio by the user, Extracting the document information of the number of extractions or the extraction ratio from among the plurality of document information in descending order of the evaluation value, Outputting, from an output unit, data in which the extracted document information is arranged in the order of the evaluation values. A control program that causes the information processing apparatus to perform the following:
13. A control program for an information processing apparatus, acquiring a plurality of document information, inputting the plurality of document information into a first pre-trained model pre-trained to output an evaluation value of a first part of the predetermined document information when the predetermined document information is input, thereby specifying the evaluation value of the first part of each document information, receiving a specification of the number of extractions or the extraction ratio by the user, extracting document information of the number of extractions or the extraction ratio from the plurality of document information in descending order of the evaluation value, inputting the extracted document information into a second pre-trained model pre-trained to output a second evaluation value of a second part of the predetermined document information when the predetermined document information is input, thereby specifying the second evaluation value of the second part of each of the extracted document information, outputting, from an output unit, data in which the extracted document information is arranged in the order of the second evaluation value, A control program that causes the information processing apparatus to perform the following: