Document Data Processing System
The document data processing system addresses the inefficiencies in document searching by using natural language queries and distributed word representations to present highly relevant document blocks, enhancing search efficiency and relevance.
Patent Information
- Application Number
- JP2024035605
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-03
- Filing Date
- 2024-03-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-09-22
AI Technical Summary
Identifying a document containing desired information from a large number of documents is labor-intensive, especially when using keyword searches or structural analysis, which can be inefficient due to numerous keywords or varied document structures.
A document data processing system and method that allows natural language input as a query sentence, divides documents into blocks, and uses distributed word representations to calculate similarity with query sentences, presenting highly relevant blocks to the user.
Enables efficient searching and presentation of highly relevant document portions to the user, reducing the burden of keyword selection and accommodating diverse document structures.
Smart Images

Figure 0007678909000003 
Figure 0007678909000004 
Figure 0007678909000005
Abstract
Description
[Technical field]
[0001] 1. Field of the Invention One aspect of the present invention relates to a document data processing method and a document data processing system, a document search method and a document search system, and a document reading support method and a document reading support system. [Background technology]
[0002] Generally, when identifying a document that is most relevant to the information a user is looking for from among a large number of documents, or when identifying a sentence or paragraph that describes that information, a search using text may be performed. Searches may also be performed using document classification information, such as the International Patent Classification in patent documents. Such searches may be used appropriately to narrow down the number of documents to a certain number, after which the contents may be examined manually. For electronic documents, there is also a method of browsing the documents while searching for keywords to find the desired information. A method of performing structure analysis of documents according to set rules has also been proposed (Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2014-219833 A [Non-patent literature]
[0004] [Non-Patent Document 1] BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Devlin et al. (Submitted on 11 Oct 2018 (v1), last revised 24 May 2019 (this version, v2)), [online], Internet<URL:https: / / arxiv.org / abs / 1810.04805v2> Summary of the Invention [Problem to be solved by the invention]
[0005] As described above, it is laborious to identify a document containing the desired information from among a number of documents narrowed down to a certain number by a primary search using keywords or classification, and to identify a highly relevant portion from among the multiple documents. In such a task, there is a method of searching the entire document for a sentence or paragraph containing a keyword by a text search using the keyword, but the desired information may not be found efficiently. Reasons for not being able to find the information efficiently include that there are too many hits for the keyword, it takes too much time to reach the desired information, and that an appropriate keyword cannot be found. In addition, when performing a structural analysis of a document according to rules, the structure to be read is limited, making it difficult to handle documents with various structures. One aspect of the present invention is to solve at least one of these problems.
[0006] One aspect of the present invention has the objective of providing a document data processing system or a document data processing method that enables input of natural language as a query sentence, enables searching of multiple documents, and presents to the reader parts of the sentence that are highly relevant to the input sentence.
[0007] Note that the description of these problems does not preclude the existence of other problems. One embodiment of the present invention does not necessarily have to solve all of these problems. Problems other than these can be extracted from the description of the specification, drawings, and claims. [Means for solving the problem]
[0008] One aspect of the present invention is a document data processing system including a document reading unit that reads a plurality of target documents, a document division unit that divides each of the plurality of target documents into a plurality of blocks, a first distributed representation acquisition unit that acquires distributed representations of words for each block, a first distributed representation holding unit that stores the distributed representations acquired by the first distributed representation acquisition unit for each target document and for each block, a query sentence reading unit that reads a query sentence, a second distributed representation acquisition unit that extracts words included in the query sentence and acquires distributed representations of the words included in the query sentence, a second distributed representation holding unit that stores the distributed representations acquired by the second distributed representation acquisition unit, and a similarity calculation unit that compares the distributed representations of the words included in the query sentence with the distributed representations of the words included in each of the plurality of blocks and calculates a similarity for each block, wherein the similarity calculation unit searches for words that match with words included in the query sentence among the words included in the blocks, and calculates a similarity between the distributed representations of the words in the block and the distributed representations of the words in the query sentence for the matching words.
[0009] One aspect of the present invention is a document data processing method including the steps of reading a plurality of target documents, dividing each of the plurality of target documents into a plurality of blocks, obtaining embedded representations of words for each block, reading a query sentence, extracting words contained in the query sentence and obtaining embedded representations of the words contained in the query sentence, and comparing the embedded representations of the words contained in the query sentence with the embedded representations of the words contained in each of the plurality of blocks and calculating a similarity for each block, wherein in the step of calculating the similarity for each block, words that match words contained in the query sentence are searched for among the words contained in the block, and for the matching words, a similarity between the embedded representations of the words in the block and the embedded representations of the words in the query sentence is calculated.
[0010] The method of displaying the score resulting from the calculation of similarity can be determined appropriately depending on the purpose of the work. For example, the text can be displayed on the screen in order of the blocks with the highest similarity. This is useful when searching for one or more documents that are most relevant among a set of multiple target documents. Alternatively, when evaluating each of a set of target documents, it is possible to display the block with the highest similarity in each of the target documents, or to display a predetermined number of blocks with the highest similarity.
[0011] Each of the multiple blocks may contain one or more paragraphs of the target document.
[0012] Each of the multiple blocks can contain one or more sentences.
[0013] The calculation of the similarity may be performed only for a predetermined part of speech.
[0014] The similarity may be calculated by calculating the cosine similarity.
[0015] If there are multiple words that match between the query sentence and the block, the sum of the similarities of the distributed representations for each word may be used as the score for that block. Effect of the Invention
[0016] One aspect of the present invention provides a document data processing method and document data processing system that enable input of natural language as a query sentence and presents to a reader parts of a plurality of documents that are highly relevant to the input sentence.
[0017] Note that the description of these effects does not preclude the existence of other effects. One embodiment of the present invention does not necessarily have all of these effects. Effects other than these can be extracted from the description in the specification, drawings, and claims. [Brief description of the drawings]
[0018] [Figure 1] FIG. 1 is a diagram illustrating an example of a document data processing system. [Diagram 2] FIG. 2 is a flowchart showing an example of a document data processing method. [Diagram 3] FIG. 3 is a flowchart showing an example of a document data processing method. [Figure 4] FIG. 4 is a diagram illustrating embedded representations of words. [Diagram 5] FIG. 5 is a diagram for explaining an example of a method for calculating the similarity. [Figure 6] FIG. 6 is a diagram illustrating an example of hardware of a document data processing system. [Figure 7] FIG. 7 is a diagram illustrating an example of hardware of a document data processing system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be easily understood by those skilled in the art that the modes and details of the present invention can be modified in various ways without departing from the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited to the description of the embodiments shown below.
[0020] In the configuration of the invention described below, the same parts or parts having similar functions are denoted by the same reference numerals in different drawings, and the repeated explanations are omitted. In addition, when referring to similar functions, the same hatch pattern may be used and no particular reference numeral may be used.
[0021] In addition, for ease of understanding, the position, size, range, etc. of each component shown in the drawings may not represent the actual position, size, range, etc. Therefore, the disclosed invention is not necessarily limited to the position, size, range, etc. disclosed in the drawings.
[0022] (Embodiment 1) In this embodiment, a document data processing system and a document data processing method according to one embodiment of the present invention will be described with reference to FIGS. 1 to 5. FIG.
[0023] In the document data processing method of this embodiment, first, a plurality of documents to be processed (target documents) are acquired. The plurality of documents are documents collected by some method, and the acquisition method is not limited to a specific method or means. For example, documents may be collected using a general search service, or documents collected by a user using a unique method. The number of target documents can be appropriately determined by the user, taking into consideration the capacity and load of the computer and memory that will perform the processing. Each target document is divided into a plurality of blocks (e.g., paragraphs), and embedded representations of words are acquired for each block. As a result, data having embedded representations of words for each block is created for each document.
[0024] On the one hand, a query sentence for obtaining information of the user's interest is obtained, and furthermore, a distributed representation of the words contained in the query sentence is obtained.
[0025] Next, words contained in the block are searched for that match a word contained in the query sentence. Then, for the matched words, the similarity (e.g., cosine similarity) between the embedded representation of the word in the block and the embedded representation of the word in the query sentence is calculated. If there are multiple matched words, the sum of the similarities of the embedded representations for each word is used as the score of the block. A block with a relatively high score is considered to have a high relevance to the query sentence. This makes it possible to identify blocks that are highly related or similar to the information from within the entire data. For example, the blocks can be arranged in descending order of score, and the blocks can be displayed on a screen used by the user in descending order of relevance.
[0026] In the document data processing method of the present embodiment, when a question sentence is input in natural language, parts related to the question sentence can be presented from among a plurality of target documents. Since different embedded representations are used for the same word depending on the sentence, blocks that are more related to or similar to the question sentence can be presented. This makes it possible to perform efficient reading and searching by, for example, creating a set by a primary search using keywords or classification, and then processing the documents included in the set. In other words, the document data processing system and document data processing method of the present embodiment can be used for document search, document reading support, and the like.
[0027] The query can contain one or more sentences. Since there is no need to select keywords to use in the search, users can search for desired information from documents with little effort.
[0028] Unless otherwise specified in this specification, a document is a description of an event in natural language that is computerized and machine-readable. Examples of documents include, but are not limited to, patent applications, court decisions, contracts, terms and conditions, product manuals, novels, publications, white papers, technical documents, etc. In this specification, a sentence includes one or more sentences.
[0029] In this specification, a word is the smallest linguistic unit that has a linguistic sound, a meaning, and a grammatical function. However, a word may be further divided into subwords to obtain a distributed representation. For example, the English word "transformer" may be divided into the subwords "transform" and "er", and a distributed representation may be given to each of them. Alternatively, a distributed representation may be given to a phrase consisting of two or more words. In this specification, a subword obtained by dividing a word is also called a word. In this specification, a phrase, word, or subword to which a distributed representation is given may also be called a token.
[0030] In this embodiment, the embedded representation of a word is obtained using a language model that can obtain different embedded representations depending on the distribution or context of surrounding words even for the same word. Alternatively, the embedded representation is obtained using a language model that can obtain different embedded representations depending on the context even for the same word. In addition, a language model that can obtain embedded representations in which the position of a word in a sentence, segment (information on the connection of a sentence), and token information are embedded may be used as the embedded representation of a word. In addition, a language model that has a self-attention function and obtains embedded representations by learning from both directions of a sentence may be used. BERT (Bidirectional Encoder Representations from Transformers) (see Non-Patent Document 1) is an example of a language model that can obtain different embedded representations depending on the distribution or context of surrounding words even for the same word.
[0031] Figure 4 plots the embedded representations obtained by BERT for "carbon" in six English sentences containing "carbon" on the XY coordinates. The three plots on the left half (squares) are sentences that contain "carbon" as an impurity in materials, and the three plots on the right half (diamonds) are sentences about "carbon" as an anode material. Figure 4 is an example showing that even for the same "carbon," different embedded representations can be obtained depending on the context and sentence.
[0032] By using a language model that can obtain different distributed representations of the same word depending on the sentence it is included in, it is possible to find blocks that are highly relevant to the information the user needs with high accuracy. For example, if the query sentence contains "carbon" as an anode material, the score of the block that contains "carbon" as an anode material will be relatively high, and the score of the block that contains "carbon" as an impurity will be relatively low.
[0033] [Document data processing system] FIG. 1 is a block diagram showing the configuration of a document data processing system 100. As shown in FIG.
[0034] Document data processing system 100 may be provided in an information processing device such as a personal computer used by a user, or a processing unit of document data processing system 100 may be provided in a server, and the system may be accessed and used from a client PC via a network.
[0035] The document data processing system 100 includes a document reading unit 101, a question input unit 102, a document division unit 103, a distributed expression acquisition unit 104a, a distributed expression acquisition unit 104b, a distributed expression storage unit 105a, a distributed expression storage unit 105b, a word selection unit 106, a similarity calculation unit 107, a score display unit 108, and a sentence display unit 109.
[0036] The document reading unit 101 reads a plurality of documents to be read.
[0037] The multiple documents read by the document reading unit 101 are a collection of documents collected by some method. For example, they may be a collection of documents collected via the Internet. They may also be documents stored in a personal computer used by a user, or documents stored in storage connected via a network.
[0038] The question input section 102 is a section where the user inputs a sentence to be specified for search.
[0039] The method of inputting a question sentence (also called a query sentence) may be to directly input any sentence, or to copy and paste text from a document file. Alternatively, a mechanism may be used in which a user arbitrarily specifies a part of a document read by the document reading unit 101 and causes the question sentence input unit 102 to read the part.
[0040] The document divider 103 divides each of the documents read by the document reader 101 into a plurality of blocks.
[0041] The document may be divided into blocks by one paragraph, one sentence separated by a period or a given number of paragraphs or sentences. Some documents have paragraph numbers included in them, and the document may be divided into blocks according to the paragraph numbers.
[0042] The embedded representation acquiring unit 104a processes each document read by the document reading unit 101 for each block, and acquires embedded representations of words included in the block.
[0043] The dispersed expression acquiring unit 104b acquires dispersed expressions of words included in the sentence input to the question input unit 102.
[0044] It is preferable that the distributed expression acquisition unit 104a and the distributed expression acquisition unit 104b basically use the same language model.
[0045] The shared representation holding unit 105a and the shared representation holding unit 105b hold the acquired shared representations as data. The data structure of the holding unit of the shared representation holding unit 105a is shown in Table 1, and the data structure of the holding unit of the shared representation holding unit 105b is shown in Table 2.
[0046] [Table 1]
[0047] [Table 2]
[0048] The shared representation holding unit 105a has a data area for each of multiple documents, and further has a data area for each block. The data area for each block stores words extracted from the sentences contained in the respective block and the shared representations corresponding to the words. Although the shared representation holding unit 105b shows a data structure assuming a single query sentence, it may also be possible to input multiple query sentences and hold the shared representations of the words in each sentence.
[0049] The word selection section 106 is a section that selects words to be used in similarity calculation from among the words contained in the input question sentence.
[0050] It is possible to select all words, select a specific part of speech such as a noun, or allow the user to freely select words. At least one word is selected, and even if only one word is selected, different embedded representations are obtained depending on the sentence and context, so scoring is possible.
[0051] The similarity calculation unit 107 calculates the similarity for the question sentence for each block by using the dispersed representations of words obtained by the dispersed representation acquisition unit 104a and the dispersed representation acquisition unit 104b.
[0052] The score display unit 108 can display the score calculated by the similarity calculation unit 107 .
[0053] The text display unit 109 can display the document read by the document reading unit 101. The text display unit 109 may further display the text inputted to the question input unit .
[0054] It is preferable that the score display unit 108 and the text display unit 109 are synchronized. For example, the display method of the target document may be changed based on the score value, such as sorting the text blocks in descending order of score, or displaying only blocks with a score equal to or greater than a predetermined value.
[0055] [Document data processing method] 2 and 3 are flowcharts illustrating the flow of processing executed by document data processing system 100. In other words, each of Fig. 2 and Fig. 3 can be said to be a flowchart illustrating an example of a document data processing method according to an aspect of the present invention.
[0056] [Step S1: Acquire multiple target documents] First, a plurality of documents to be interpreted are read by document reading unit 101 of document data processing system 100 .
[0057] [Step S2: Divide the target document into multiple blocks] Next, the document division unit 103 divides each of the multiple target documents into multiple blocks.
[0058] [Step S3: Obtain distributed representations of words for each block] Next, the sentence is input to the dispersed representation acquiring unit 104a for each block, and dispersed representations of words are acquired. Specifically, the target document is input to a language model such as BERT for each block, and dispersed representations of words are acquired. The dispersed representations acquired by the dispersed representation acquiring unit 104a are stored in the dispersed representation holding unit 105a for each target document and for each block.
[0059] [Step S4: Obtain the query sentence] Furthermore, a query sentence is acquired by the question sentence input unit 102 of the document data processing system 100. The query sentence may be a sentence arbitrarily input by the user, or may be a sentence of a part of the target document that is of high interest to the user. In FIG. 2, an example is shown in which steps S4 and S5 are performed after step S3, but as shown in FIG. 3, steps S1 to S3 and steps S4 and S5 can be performed independently of each other, and the order does not matter.
[0060] [Step S5: Obtain distributed representations of words contained in the query sentence] Next, the query sentence is input to the distributed representation acquiring unit 104b, and distributed representations of the words are acquired. Specifically, the query sentence is input to a language model such as BERT, and distributed representations of the words are acquired. The distributed representations acquired by the distributed representation acquiring unit 104b are stored in the distributed representation holding unit 105b.
[0061] [Step S6: Calculate the score of the block] Next, the similarity calculation unit 107 searches for matching words between the words contained in each block and the words contained in the query sentence, and only if the words match, calculates the cosine similarity between the embedded representations of the matching words, and obtains the score of the block by calculating the sum of the cosine similarities within the block.
[0062] The word selection unit 106 may select words to be used in the similarity calculation from among the words included in the query sentence, and perform the similarity calculation only for the selected words.
[0063] In this embodiment, an example in which the similarity is calculated using cosine similarity is shown, but other similarity calculation methods may be used.
[0064] A method of calculating the score for each block will be described with reference to FIG. 5. FIG. 5 shows an example of comparing blocks 1, 2, 3, and 4 of target document 1 and target document 2 with the query sentence. First, in each block of the target document, a word that matches a word in the query sentence is searched for, and the cosine similarity of the distributed representation of the word is calculated only for the matched word. If there are multiple matched words in one block, the cosine similarity of each word is added to calculate the score of the block. For example, in block 1 of target document 1 shown in FIG. 5, two words, word W1 and word W2, match in the query sentence. In this case, the score of block 1 of target document 1 is the sum of the cosine similarity of word W1 and the cosine similarity of word W2.
[0065] [Step S7: Output the calculated score] Then, blocks with high calculated scores can be presented to the user as blocks that are likely to contain the desired information. Examples of the presentation method include a method of setting a predetermined threshold and presenting blocks that exceed the threshold, a method of presenting the block with the maximum score in each document, or a method of presenting a predetermined number of blocks with the highest scores among all the multiple blocks. These methods may also be combined as appropriate.
[0066] As described above, in the document data processing system and document data processing method of this embodiment, when a user provides a set of documents that the user wishes to read and understand, and sentences related to the information the user needs, the system can present blocks in the set of documents that are highly relevant to the information the user needs. This eliminates the need for the user to select keywords, making it easier for the user to find the desired information from the documents.
[0067] The document data processing system and document data processing method of the present embodiment use a language model that can obtain different distributed representations of the same word depending on the sentence it is included in. This makes it possible to find blocks that are highly relevant to the information the user needs with high accuracy.
[0068] This embodiment mode can be combined with other embodiment modes as appropriate. In addition, in the case where a plurality of configuration examples are shown in one embodiment mode in this specification, the configuration examples can be combined as appropriate.
[0069] (Embodiment 2) In this embodiment, a document data processing system according to one embodiment of the present invention will be described with reference to FIGS.
[0070] The document data processing system of this embodiment can easily search for and acquire desired information from a document by using the document data processing method shown in the first embodiment.
[0071] <Document data processing system configuration example 1> 6 shows a block diagram of document data processing system 200. In the drawings attached to this specification, the components are classified by function and shown as independent blocks in the block diagram, but in reality, it is difficult to completely separate the components by function, and one component may be involved in multiple functions. Also, one function may be involved in multiple components, and for example, the processing performed by processing unit 120 may be executed by different servers depending on the processing.
[0072] Document data processing system 200 has at least a processing unit 120. Document data processing system 200 shown in FIG.
[0073] [Input section 110] A question sentence (query sentence) is supplied to the input unit 110 from outside the document data processing system 200. A set of target documents may also be supplied to the input unit 110 from outside the document data processing system 200. The set of target documents and the query sentence supplied to the input unit 110 are each supplied to the processing unit 120, the storage unit 130, or the database 140 via a transmission path 160.
[0074] The target document and the query sentence are input as, for example, text data, audio data, or image data. The target document is preferably input as text data.
[0075] Methods for inputting a query sentence include, for example, key input using a keyboard, a touch panel, etc., voice input using a microphone, reading from a recording medium, image input using a scanner, a camera, etc., and acquisition via communication.
[0076] The document data processing system 200 may have a function of converting voice data into text data. For example, the processing unit 120 may have this function. Alternatively, the document data processing system 200 may further have a voice conversion unit having this function.
[0077] The document data processing system 200 may have an optical character recognition (OCR) function. This allows characters included in image data to be recognized and text data to be created. For example, the processing unit 120 may have this function. Alternatively, the document data processing system 200 may further have a character recognition unit having this function.
[0078] [Processing section 120] The processing unit 120 has a function of performing calculations using data supplied from the input unit 110, the storage unit 130, the database 140, etc. The processing unit 120 can supply the calculation results to the storage unit 130, the database 140, the display unit 150, etc.
[0079] The processing unit 120 has a function of dividing a document into a plurality of blocks. For example, the processing unit 120 may have a function of dividing a document into a plurality of blocks, such as by chapter, by paragraph, or by a predetermined number of sentences.
[0080] The processing unit 120 has a function of acquiring embedded representations of words. For example, it is possible to acquire embedded representations of words included in a block of a target document or words included in a query sentence.
[0081] The processing unit 120 has a function of extracting words from a query sentence, which makes it possible to select words to be used in the similarity calculation from among the words contained in the query sentence.
[0082] The processing unit 120 has a function of calculating the similarity between embedded representations of words.
[0083] The processing unit 120 may be a transistor having a metal oxide in a channel formation region. Since the off-state current of the transistor is extremely low, the transistor can be used as a switch for holding charge (data) flowing into a capacitive element functioning as a memory element, thereby ensuring a long data holding period. By using this characteristic in at least one of the register and cache memory of the processing unit 120, the processing unit 120 can be operated only when necessary, and in other cases, the processing unit 120 can be turned off by saving information of the immediately preceding process in the memory element. In other words, normally-off computing is possible, and the power consumption of the document data processing system can be reduced.
[0084] Note that in this specification and the like, a transistor using an oxide semiconductor for a channel formation region is referred to as an oxide semiconductor transistor or an OS transistor. The channel formation region of an OS transistor preferably contains a metal oxide.
[0085] The metal oxide in the channel formation region preferably contains indium (In). When the metal oxide in the channel formation region contains indium, the carrier mobility (electron mobility) of the OS transistor is increased. The metal oxide in the channel formation region is preferably an oxide semiconductor containing an element M. The element M is preferably aluminum (Al), gallium (Ga), or tin (Sn). Other elements that can be used for the element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, the element M may be a combination of a plurality of the above elements. The element M is, for example, an element having a high bond energy with oxygen. For example, it is an element that has a higher bond energy with oxygen than indium. In addition, the metal oxide in the channel formation region preferably contains zinc (Zn). Metal oxides containing zinc may be easily crystallized.
[0086] The metal oxide in the channel formation region is not limited to a metal oxide containing indium. The semiconductor layer may be a metal oxide containing zinc but not indium, such as zinc tin oxide or gallium tin oxide, a metal oxide containing gallium, or a metal oxide containing tin.
[0087] The processing section 120 may also use a transistor containing silicon in the channel formation region.
[0088] The processing section 120 may include a combination of a transistor including an oxide semiconductor in a channel formation region and a transistor including silicon in a channel formation region.
[0089] The processing unit 120 includes, for example, an arithmetic circuit or a central processing unit (CPU: Central Processing Unit).
[0090] The processing unit 120 may have a microprocessor such as a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The microprocessor may be realized by a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The processing unit 120 can perform various data processing and program control by interpreting and executing instructions from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the memory area of the processor and the storage unit 130.
[0091] The processing unit 120 may include a main memory. The main memory includes at least one of a volatile memory such as a RAM and a non-volatile memory such as a ROM.
[0092] The RAM may be, for example, a dynamic random access memory (DRAM) or a static random access memory (SRAM), and a virtual memory space is allocated and used as a working space for the processing unit 120. The operating system, application programs, program modules, program data, lookup tables, and the like stored in the storage unit 130 are loaded into the RAM for execution. The data, programs, and program modules loaded into the RAM are each directly accessed and operated by the processing unit 120.
[0093] ROM can store BIOS (Basic Input / Output System) and firmware that do not require rewriting. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows the erasure of stored data by exposure to ultraviolet light, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.
[0094] [Storage section 130] The storage unit 130 has a function of storing a program executed by the processing unit 120. The storage unit 130 may also have a function of storing, for example, a calculation result generated by the processing unit 120 and data input to the input unit 110. Specifically, the storage unit 130 preferably has a function of storing embedded representations of words acquired by the processing unit 120.
[0095] The storage unit 130 has at least one of a volatile memory and a non-volatile memory. The storage unit 130 may have a volatile memory such as a DRAM or an SRAM. The storage unit 130 may have a non-volatile memory such as a ReRAM (Resistive Random Access Memory, also called a resistance change type memory), a PRAM (Phase change Random Access Memory), a FeRAM (Ferroelectric Random Access Memory), an MRAM (Magnetoresistive Random Access Memory, also called a magnetoresistive type memory), or a flash memory. The storage unit 130 may also have a recording media drive such as a hard disk drive (HDD) and a solid state drive (SSD).
[0096] [Database 140] The document data processing system may include a database 140. For example, the database 140 has a function of storing a plurality of documents. For example, the document data processing method according to one aspect of the present invention may be applied to a collection of documents stored in the database 140. Note that the storage unit 130 and the database 140 do not need to be separated from each other. For example, the document data processing system may include a storage unit having the functions of both the storage unit 130 and the database 140.
[0097] The memories included in the processing unit 120, the storage unit 130, and the database 140 can each be considered as an example of a non-transitory computer-readable storage medium.
[0098] [Display section 150] The display unit 150 has a function of displaying the calculation results from the processing unit 120. The display unit 150 also has a function of displaying a target document. The display unit 150 may also have a function of displaying a query sentence.
[0099] The document data processing system 200 may have an output unit. The output unit has a function of supplying data to an external device.
[0100] [Transmission Line 160] The transmission path 160 has a function of transmitting various data. Data can be transmitted and received between the input unit 110, the processing unit 120, the storage unit 130, the database 140, and the display unit 150 via the transmission path 160. For example, data such as a target document is transmitted and received via the transmission path 160.
[0101] <Document data processing system configuration example 2> 7 shows a block diagram of document data processing system 210. Document data processing system 210 includes server 220 and terminal 230 (such as a personal computer).
[0102] The server 220 includes a communication unit 161a, a transmission path 162, a processing unit 120, and a storage unit 170. Although not shown in FIG. 7, the server 220 may further include an input / output unit and the like.
[0103] The terminal 230 includes a communication unit 161b, a transmission path 164, a processing unit 180, a storage unit 130, and a display unit 150. Although not shown in FIG 7, the terminal 230 may further include a database and the like.
[0104] A user of document data processing system 210 inputs a question (query sentence) into input unit 110 of terminal 230. The question sentence is transmitted from communication unit 161b of terminal 230 to communication unit 161a of server 220.
[0105] The question received by the communication unit 161a is stored in the storage unit 170 via the transmission path 162. Alternatively, the question may be supplied to the processing unit 120 directly from the communication unit 161a.
[0106] The document division, distributed representation acquisition, and similarity calculation described in the first embodiment each require high processing power. The processing unit 120 of the server 220 has a higher processing power than the processing unit 180 of the terminal 230. Therefore, it is preferable that each of these processes is performed by the processing unit 120.
[0107] Then, the processing unit 120 calculates a score for the block. The score is stored in the storage unit 170 via the transmission path 162. Alternatively, the score may be directly supplied from the processing unit 120 to the communication unit 161a. The score is transmitted from the communication unit 161a of the server 220 to the communication unit 161b of the terminal 230. The score is displayed on the display unit 150 of the terminal 230.
[0108] [Transmission path 162 and transmission path 164] The transmission paths 162 and 164 have a function of transmitting data. Data can be transmitted and received between the communication unit 161a, the processing unit 120, and the storage unit 170 via the transmission path 162. Data can be transmitted and received between the input unit 110, the communication unit 161b, the processing unit 180, the storage unit 130, and the display unit 150 via the transmission path 164.
[0109] [Processing section 120 and processing section 180] The processing unit 120 has a function of performing calculations using data supplied from the communication unit 161a, the storage unit 170, etc. The processing unit 180 has a function of performing calculations using data supplied from the communication unit 161b, the storage unit 130, the display unit 150, etc. For the processing unit 120 and the processing unit 180, refer to the description of the processing unit 120. It is preferable that the processing unit 120 has a higher processing capacity than the processing unit 180.
[0110] [Storage section 130] The storage unit 130 has a function of storing a program executed by the processing unit 180. The storage unit 130 also has a function of storing the calculation result generated by the processing unit 180, the data input to the communication unit 161b, the data input to the input unit 110, and the like.
[0111] [Storage section 170] The storage unit 170 has a function of storing a plurality of documents, the calculation results generated by the processing unit 120, data input to the communication unit 161a, and the like.
[0112] [Communication Unit 161a and Communication Unit 161b] Using the communication units 161a and 161b, data can be transmitted and received between the server 220 and the terminal 230. A hub, a router, a modem, or the like can be used as the communication units 161a and 161b. Data can be transmitted and received using either a wired connection or wirelessly (for example, radio waves, infrared rays, etc.).
[0113] In addition, communication between server 220 and terminal 230 may be performed by connecting to a computer network such as the Internet, which is the basis of the World Wide Web (WWW), an intranet, an extranet, a Personal Area Network (PAN), a Local Area Network (LAN), a Campus Area Network (CAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), or a Global Area Network (GAN).
[0114] This embodiment mode can be combined with other embodiment modes as appropriate. [Explanation of symbols]
[0115] W1: word, W2: word, 1: block, 2: block, 3: block, 4: block, 100: document data processing system, 101: document reading unit, 102: question sentence input unit, 103: document division unit, 104a: distributed representation acquisition unit, 104b: distributed representation acquisition unit, 105a: distributed representation holding unit, 105b: distributed representation holding unit, 106: word selection unit, 107: similarity calculation unit, 108: score display unit, 109: sentence display unit, 110: input unit, 120: processing unit, 130: memory unit, 140: database, 150: display unit, 160: transmission path, 161a: communication unit, 161b: communication unit, 162: transmission path, 164: transmission path, 170: memory unit, 180: processing unit, 200: document data processing system, 210: document data processing system, 220: server, 230: terminal
Claims
1. a document reading unit, a document dividing unit, a first distributed representation acquisition unit, a first distributed representation holding unit, a query sentence reading unit, a second distributed representation acquisition unit, a second distributed representation holding unit, a similarity calculation unit, a score display unit, and a sentence display unit, The document reading unit has a function of reading a plurality of target documents, the document division unit has a function of dividing each of the plurality of target documents into a plurality of blocks; the first distributed representation acquisition unit has a function of acquiring distributed representations of words for each of the blocks; the first distributed representation holding unit has a function of storing the distributed representations acquired by the first distributed representation acquisition unit for each of the target documents and for each of the blocks; the query sentence reading unit has a function of reading a query sentence, the second distributed representation acquisition unit has a function of extracting words included in the query sentence and acquiring distributed representations of the words included in the query sentence; the second distributed representation holding unit has a function of storing the distributed representation acquired by the second distributed representation acquisition unit, the similarity calculation unit has a function of comparing a distributed representation of a word included in the query sentence with a distributed representation of a word included in each of the plurality of blocks, and calculating a similarity for each of the blocks; and when there are a plurality of words that match between the query sentence and the block, a function of calculating a sum of similarities of the distributed representations for each of the words as a score of the block; The score display unit has a function of displaying the score, the sentence display unit has a function of sorting and displaying the plurality of blocks in descending order of the score, the similarity calculation unit searches for words that match with words included in the query sentence from among the words included in the block, and calculates a similarity between a distributed representation of the word in the block and a distributed representation of the word in the query sentence for the matching word; Document data processing system.
2. In claim 1, each of the plurality of blocks includes one or more paragraphs of the target document; Document data processing system.
3. In claim 1, each of the plurality of blocks includes one or more sentences; Document data processing system.
Citation Information
Patent Citations
Document reading comprehension support device, document reading comprehension support system, and program
JP2014219833A
Retrieval device, similarity calculation method, and program
JP2019082931A
Document search using grammatical units
US20190155913A1