Document data processing system

The document data processing system efficiently retrieves relevant document sections by using natural language input and distributed representation analysis, addressing inefficiencies in existing search methods.

JP2025111781AInactive Publication Date: 2025-07-30SEMICON ENERGY LAB CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025076557
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2025-05-02
Publication Date
2025-07-30
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure 2025111781000001_ABST
    Figure 2025111781000001_ABST
Patent Text Reader

Abstract

To provide a document data processing system configured to enable input of natural language as query text and searching a plurality of documents, and present a portion highly relevant to the input text to a reader.SOLUTION: A document data processing system 100 includes: a document read unit 101 that reads a plurality of target documents; a document division unit 103 that divides each of the plurality of target documents into a plurality of blocks; a distributed representation acquisition unit 104a that acquires a distributed representation of a word in each of the blocks; a distributed representation retention unit 105a that stores the acquired distributed representation, by target document and by block; a query text input unit 102 that inputs a query text; a distributed representation acquisition unit 104b that extracts a word included in the query text and acquires a distributed representation of the word; a distributed representation retention unit 105b that stores the acquired distributed representation; and a similarity calculation unit 107 that compares the distributed representation of the word included in the query text with the distributed representation of the word included in each of the blocks to calculate similarity by block.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a method for processing document data and a document data processing system. One aspect of the present invention relates to a document search method, a document search system, a document reading support method, and a document reading support system.

Background Art

[0002] Generally, when identifying a document most relevant to information required by a user from a large number of documents, or when identifying a sentence or paragraph describing that information, text-based search may be performed. Also, search may be performed using classification information of documents such as the International Patent Classification in patent documents. After narrowing down the number of documents to a certain number by appropriately using such search and then manually examining the contents, in the case of electronic documents, there is also a method of browsing the documents while performing a search using words as keywords to find desired information. Further, a method of performing structural analysis of documents according to set rules has been proposed (Patent Document 1). (Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Selecting a document that describes the target information from a plurality of documents narrowed down to a certain number by primary search using keywords and classification as described above, and determining the relevance from a plurality of documents Identifying high points is a labor-intensive task. In such a task, there is also a method of searching for sentences or paragraphs containing keywords from the entire document by text search using keywords However, there are cases where the desired information cannot be found efficiently. Reasons for inefficient finding include that there are too many hits with keywords and it takes too much time to reach the desired information, and that appropriate keywords cannot be found. Also, when performing structural analysis of a document according to rules Since the structure to be read is limited, it is difficult to handle documents with various structures. One aspect of the present invention solves at least one of these problems That is, it takes too long to reach the desired information because there are too many hits with keywords, and appropriate keywords cannot be found. Also, when performing structural analysis of a document according to rules Since the structure to be read is limited, it is difficult to handle documents with various structures. One aspect of the present invention solves at least one of these problems That is, it takes too long to reach the desired information because there are too many hits with keywords, and appropriate keywords cannot be found. Also, when performing structural analysis of a document according to rules Since the structure to be read is limited, it is difficult to handle documents with various structures. One aspect of the present invention solves at least one of these problems That is, it is.

[0006] One aspect of the present invention enables natural language input as a query sentence, enables retrieval of a plurality of documents, and presents to the reader a document data processing system that presents a portion highly relevant to the input sentence to the reader. One of the problems is to provide a method for processing stem or document data.

[0007] Note that the description of these problems does not prevent the existence of other problems. One aspect of the present invention does not necessarily need to solve all of these problems. It is possible to extract other problems from the descriptions in the specification, drawings, and claims.

Means for Solving the Problems

[0008] One aspect of the present invention includes a document reading unit that reads a plurality of target documents, a document splitting unit that splits each of the plurality of target documents into a plurality of blocks, a first dispersion representation acquisition unit that acquires a dispersion representation of words for each block, a first dispersion representation holding unit that stores the dispersion representation acquired by the first dispersion representation acquisition unit for each target document and for each block, a query sentence reading unit that reads a query sentence, a second dispersion representation acquisition unit that extracts words included in the query sentence and acquires a dispersion representation of the words included in the query sentence, a second dispersion representation holding unit that stores the dispersion representation acquired by the second dispersion representation acquisition unit, and a similarity calculation unit that compares the dispersion representation of the words included in the query sentence with the dispersion representation of the words included in each of the plurality of blocks and calculates the similarity for each block. The similarity calculation unit searches for words that match the words included in the query sentence from among the words included in the block, and calculates the similarity between the dispersion representation of the words in the block and the dispersion representation of the words in the query sentence.

[0009] One aspect of the present invention includes a step of reading a plurality of target documents, and a step of splitting each of the plurality of target documents into a plurality of The step of dividing into a number of blocks, the step of obtaining the distributed representation of words for each block, the step of reading the query sentence, extracting the words included in the query sentence, and obtaining the distributed representation of the words included in the query sentence, and comparing the distributed representation of the words included in the query sentence with the distributed representation of the words included in each of the plurality of blocks, and calculating the similarity for each block, wherein, in the step of calculating the similarity for each block, words that match the words included in the query sentence are searched for among the words included in the block, and for the matching words, the similarity between the distributed representation of the words in the block and the distributed representation of the words in the query sentence is calculated, which is a document data processing method. The method for displaying the score of the calculated similarity result can be appropriately determined according to the purpose of the operation. For example, the sentences can be displayed on the screen in the order of the blocks with higher similarity. This is useful when wanting to find one or more most relevant documents from among the entire plurality of target documents. Or, when wanting to evaluate each of the plurality of target documents, it is also possible to display the block with the highest similarity in each of the target documents, or display the top predetermined number of blocks with higher similarity.

[0010] The plurality of blocks may each include one or more paragraphs of the target document. The plurality of blocks may each include one or more sentences. The calculation of similarity may be performed only for a predetermined part of speech.

[0011]

[0012]

[0013]

[0014] ​​​​​​The similarity may be calculated by calculating the cosine similarity.

[0015] When there are multiple words that match between the query sentence and the block, the dispersion table of the similarity for each word may be used as the score of the block.

Advantages of the Invention

[0016] According to one aspect of the present invention, it is possible to input natural language as a query sentence, and among a plurality of documents, a method for processing document data and a document data processing system that present a part highly relevant to the input sentence to the reader can be provided.

[0017] Note that the description of these effects does not prevent the existence of other effects. One aspect of the present invention does not necessarily have to have all of these effects. From the descriptions of the specification, drawings, and claims, it is possible to extract effects other than these.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0019] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and those skilled in the art can easily understand that the form and details can be variously changed without departing from the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited to the description of the embodiments shown below. In the configuration of the invention described below, the same reference numerals are commonly used among different drawings for the same part or parts having the same or similar functions, and the repeated description thereof will be omitted. Also, when referring to similar functions, the hatching patterns may be the same and may not be particularly labeled. In addition, in the drawings, the positions, sizes, ranges, etc. of the respective configurations shown may not represent the actual positions, sizes, ranges, etc. for the sake of simplicity of understanding. For this reason, the disclosed invention is not necessarily limited to the positions, sizes, ranges, etc. disclosed in the drawings.

[0020] In the configuration of the invention described below, the same reference numerals are commonly used among different drawings for the same part or parts having the same or similar functions, and the repeated description thereof will be omitted. Also, when referring to similar functions, the hatching patterns may be the same and may not be particularly labeled. In addition, in the drawings, the positions, sizes, ranges, etc. of the respective configurations shown may not represent the actual positions, sizes, ranges, etc. for the sake of simplicity of understanding. For this reason, the disclosed invention is not necessarily limited to the positions, sizes, ranges, etc. disclosed in the drawings.

[0021] In addition, in the drawings, the positions, sizes, ranges, etc. of the respective configurations shown may not represent the actual positions, sizes, ranges, etc. for the sake of simplicity of understanding. For this reason, the disclosed invention is not necessarily limited to the positions, sizes, ranges, etc. disclosed in the drawings.

[0022] (Embodiment 1) In this embodiment, a document data processing system and a document data processing method according to an aspect of the present invention will be described with reference to FIGS. 1 to 5.

[0023] In the document data processing method of this embodiment, first, a plurality of documents (target documents) to be processed are acquired. The plurality of documents are documents collected by some method, and the acquisition method is not limited to a specific method or means. For example, documents collected using a general search service may be used, or documents collected by the user in an original way may be targeted. Also, the number of target documents is considered in view of the capacity and load of the computer and memory for performing the processing, and is not limited to a specific method or means. For example, documents collected using a general search service may be used, or documents collected by the user in an original way may be targeted. Also, the number of target documents is considered in view of the capacity and load of the computer and memory for performing the processing, and is not limited to a specific method or means. For example, documents collected using a general search service may be used, or documents collected by the user in an original way may be targeted. Also, the number of target documents is considered in view of the capacity and load of the computer and memory for performing the processing, and is not limited to a specific method or means. For example, documents collected using a general search service may be used, or documents collected by the user in an original way may be targeted. Also, the number of target documents is considered in view of the capacity and load of the computer and memory for performing the processing, It can be determined by the user as appropriate. Each of the target documents has a plurality of blocks (for example, paragraphs). They are delimited by blocks, and a distributed representation of words is obtained for each block. As a result, data with the distributed representation of words for each block is created for each document.

[0024] On the other hand, a query sentence for obtaining information of interest to the user is acquired, and further, a distributed representation of words included in the query sentence is obtained.

[0025] Next, words that match the words included in the query sentence are searched for among the words included in the block. Then, for the matching words, the similarity (for example, cosine similarity) between the distributed representation of the words in the block and the distributed representation of the words in the query sentence is calculated. When there are a plurality of matching words, the sum of the similarities of the distributed representations for each word is taken as the score of the block. A block with a relatively high score is considered to have a high relevance to the query sentence. Thus, a block with a high relevance or similarity to the information can be specified from the entire data. For example, the blocks can be arranged in descending order of score, and the blocks with high relevance can be displayed on the screen for the user to use.

[0026] In the document data processing method of the present embodiment, by inputting a question sentence in natural language, relevant parts to the question sentence can be presented from a plurality of target documents. Since different distributed representations are used depending on the sentence even for the same word, blocks with higher relevance or similarity to the question sentence can be presented. Thus, for example, after creating a set by primary search using keywords or classification, efficient reading can be achieved by processing the documents included in the set. ​​​​​​​​​​​​In other words, the document data processing system and The document data processing method can be used for document search, document reading support, and the like.

[0027] A query can contain one or more sentences. Therefore, the user can easily find the desired information from the document.

[0028] Unless otherwise specified in this specification, a document is a description of an event in natural language, and Documents are, for example, patent applications, legal precedents, contracts, terms and conditions, product This includes, but is not limited to, manuals, novels, publications, white papers, technical documents, etc.; and In this specification and the like, a sentence includes one or more sentences.

[0029] In this specification, a word is the smallest linguistic unit that has a linguistic sound, meaning, and grammatical function. However, it is also possible to find distributed representations for subwords that are further divided into words. The word "transformer" is a combination of the words "transform" and "er." It is also possible to decompose it into blocks and give each block a distributed representation. It is also possible to provide distributed representations for phrases that are connected by words. In this specification, the distributed representation is also called a word. A given phrase, word or subword may also be called a token.

[0030] In this embodiment, the distributed representation of a word is based on the distribution of surrounding words, even if the word is the same. Or, it is obtained using a language model that can obtain different distributed representations depending on the context. It is obtained using a language model that can obtain different distributed representations even for a single word depending on the context. Also as the distributed representation of a word, a distributed representation in which the position of the word and the segment (information on the connection of sentences) in the sentence and the information of the token are embedded may be used. Also, a language model that has a self attention function and learns from both directions of the sentence to obtain a distributed representation may be used. Even for the same word, different distributed representations are obtained depending on the distribution or context of the surrounding words. As an example of a language model BERT (Bidirectional Encoder Representations from Transformers) (see Non-Patent Document 1) can be cited. (see Non-Patent Document 1).

[0031] FIG. 4 shows the distributed representations obtained by BERT for the word "carbon" in six English sentences containing "carbon", plotted on the XY coordinates for each sentence . The three plots (squares) on the left half are sentences in which "carbon" is included as an impurity in the material, and the three plots (diamonds) on the right half are sentences about "carbon" as a negative electrode material . FIG. 4 is an example showing that different distributed representations are obtained for the same "carbon" depending on the context and the sentence . . .''

[0032] By using a language model in which different distributed representations of words are obtained depending on the sentences containing the same word, blocks highly relevant to the information required by the user can be searched for with high accuracy . For example, if the query sentence contains "carbon" as a negative electrode material, the score of the block containing "carbon" as a negative electrode material becomes relatively high . It is considered that the score of the block that becomes relatively high and contains "carbon" as an impurity will become relatively low. It is considered that it will become relatively low.

[0033] [Document Data Processing System] FIG. 1 is a block diagram showing the configuration of a document data processing system 100.

[0034] The document data processing system 100 may be provided in an information processing device such as a personal computer used by a user. Alternatively, a processing unit of the document data processing system 100 may be provided in a server, and it may be configured to be accessed and used via a network from a client PC. The document data processing system 100 may be provided in an information processing device such as a personal computer used by a user. Alternatively, a processing unit of the document data processing system 100 may be provided in a server, and it may be configured to be accessed and used via a network from a client PC. The document data processing system 100 may be provided in an information processing device such as a personal computer used by a user. Alternatively, a processing unit of the document data processing system 100 may be provided in a server, and it may be configured to be accessed and used via a network from a client PC. It may be configured to be accessed and used via a network from a client PC.

[0035] The document data processing system 100 includes a document reading unit 101, a question sentence input unit 102, a document splitting unit 103, a distributed representation acquisition unit 104a, a distributed representation acquisition unit 104b, a distributed representation storage unit 105a, a distributed representation storage unit 105b, a word selection unit 106, a similarity calculation unit 107, a score display unit 108, a distributed representation storage unit 105b, a word selection unit 106, a similarity calculation unit 107, a score display unit 108, and a sentence display unit 109.

[0036] The document reading unit 101 reads a plurality of documents to be read.

[0037] The plurality of documents read by the document reading unit 101 is a set of documents collected by some method. For example, it may be a set of documents collected via the Internet. Also, it may be a document stored in a personal computer used by a user, or a document stored in a storage connected via a network For example, it may be a set of documents collected via the Internet. Also, it may be a document stored in a personal computer used by a user, or a document stored in a storage connected via a network It may be a document stored in a storage connected via a network.

[0038] The question sentence input unit 102 is a part for a user to input a sentence specified for search.

[0039] As an input method for a question sentence (also referred to as a query sentence), an arbitrary sentence can be directly input, or text copied from a document file may be pasted. Also, a mechanism may be provided where a user can arbitrarily specify a part of the document read by the document reading unit 101 and load it into the question sentence input unit 102.

[0040] The document splitting unit 103 splits each of the plurality of documents read by the document reading unit 101 into a plurality of blocks.

[0041] It may be split with one paragraph as one block, one sentence separated by a period or full stop as one block, or a predetermined number of paragraphs or a predetermined number of sentences as one block. Depending on the document, there may be a document in a format that includes paragraph numbers from the beginning, and it may be split into blocks according to those paragraph numbers.

[0042] The distributed representation acquisition unit 104a processes each document read by the document reading unit 101 block by block and acquires the distributed representation of the words included in the block.

[0043] The distributed representation acquisition unit 104b acquires the distributed representation of the words included in the sentence input to the question sentence input unit 102.

[0044] It is preferable that the distributed representation acquisition unit 104a and the distributed representation acquisition unit 104b basically use the same language model.

[0045] The distributed representation holding unit 105a and the distributed representation holding unit 105b hold the acquired distributed representation as data. The data structure of the holding unit of the distributed representation holding unit 105a is shown in Table 1, and the data structure of the holding unit of the distributed representation holding unit 105b is shown in Table 2.

[0046] ​​​​​

Table 1

[0047]

Table 2

[0048] In the distributed representation holding unit 105a, there is a data area for each of a plurality of documents, and there is further a data area for each block. In the data area of each block, words extracted from the sentences of each block and the distributed representations corresponding to the words are stored. The distributed representation holding unit 105b shows a data structure assuming a single query sentence, but a plurality of query sentences may be input and the distributed representations of the words in each of the sentences may be held. The word selection unit 106 is a part that selects words to be used for calculating similarity among the words included in the input question sentence. All words may be selected, predetermined parts of speech such as nouns may be selected, or the user may be allowed to freely select words. At least one word is selected, and even in the case of one word, since different distributed representations can be obtained depending on the sentence or context, scoring is possible. The similarity calculation unit 107 calculates the similarity to the question sentence for each block using the distributed representations of the words obtained by the distributed representation acquisition unit 104a and the distributed representation acquisition unit 104b.

[0049]

[0050]

[0051]

[0052] The score display unit 108 can display the score calculated by the similarity calculation unit 107.

[0053] The document display unit 109 can display the document read by the document reading unit 101. The display unit 109 may further display the text input to the question text input unit 102.

[0054] The score display unit 108 and the text display unit 109 are preferably synchronized. For example, rearranging the text blocks in descending order of score, or displaying only the blocks with a score equal to or higher than a predetermined value, etc., the display method of the target document may be changed based on the score value.

[0055] [Document Data Processing Method] FIG. 2 and FIG. 3 are respectively flowcharts for explaining the processing flow executed by the document data processing system 100. That is, FIG. 2 and FIG. 3 can also be said to be flowcharts showing examples of the document data processing method according to one aspect of the present invention.

[0056] [Step S1: Obtain a plurality of target documents] First, a plurality of documents to be read are read by the document reading unit 101 of the document data processing system 100.

[0057] [Step S2: Divide the target document into a plurality of blocks] Next, the document splitting unit 103 splits each of the plurality of target documents into a plurality of blocks.

[0058] [Step S3: Obtain the distributed representation of words for each block] Next, the distributed representation acquisition unit 104a inputs the text for each block to obtain the distributed representation of words. Specifically, the target document is input to a language model such as BERT for each block to obtain the distributed representation of words. The distributed representation obtained by the distributed representation acquisition unit 104a is stored in the distributed representation storage unit 105a for each target document and for each block.

[0059] [Step S4: Obtain the query text]​​ Furthermore, the query sentence is acquired by the query sentence input unit 102 of the document data processing system 100. The query sentence may be a sentence arbitrarily input by the user, or may be a sentence in a portion where the user's interest in the target document is high. In FIG. 2, an example of performing step S4 and step S5 after step S3 is shown. However, as shown in FIG. 3, step S1 to step S3 and step S4 and step S5 can be performed independently of each other, and the order does not matter.

[0060] [Step S5: Obtain the distributed representation of the words included in the query sentence] Next, the query sentence is input to the distributed representation acquisition unit 104b to obtain the distributed representation of the words. Specifically, the query sentence is input to a language model such as BERT to obtain the distributed representation of the words. The distributed representation obtained by the distributed representation acquisition unit 104b is stored in the distributed representation holding unit 105b.

[0061] [Step S6: Calculate the score of the block] Next, the similarity calculation unit 107 searches for words that match between the words included in each block and the words included in the query sentence. Only when the words match, the cosine similarity is calculated between the distributed representations of the matching words, and the score of the block is obtained by calculating the sum of the cosine similarities within the block.

[0062] The word selection unit 106 selects the words to be used for similarity calculation from among the words included in the query sentence, and the similarity may be calculated only for the selected words.

[0063] In this embodiment, an example of calculating the similarity using the cosine similarity is shown, but other similarity calculation methods may be used.

[0064] A method of calculating a score for each block will be described with reference to FIG. 5. In FIG. 5, an example of comparing block 1, block 2, block 3, and block 4 of target document 1 and target document 2 with respect to a query sentence is shown. First, in each block of the target document, words that match the words of the query sentence are searched, and for only the matching words, the cosine similarity of the distributed representation of the words is calculated. When there are multiple matching words in one block, the score of the block is calculated by adding the cosine similarities for each word. For example, in block 1 of target document 1 shown in FIG. 5, two words, word W1 and word W2, of the query sentence match. In this case, the score of block 1 of target document 1 is the sum of the cosine similarity of word W1 and the cosine similarity of word W2.

[0065] [Step S7: Output the calculated score] Then, the block with the highest calculated score can be presented to the user as a block that is likely to contain the information being sought. As a method of presentation, a method of setting a predetermined threshold and presenting blocks that exceed the threshold, a method of presenting the block having the maximum value of the score in each document, or a method of presenting a predetermined number of blocks having the top scores among all the blocks can be mentioned. Also, these methods may be appropriately combined.

[0066] As described above, in the document data processing system and the document data processing method of the present embodiment, a set of documents that the user wants to read and a sentence related to the information needed are supplied from the user, and blocks highly relevant to the information required by the user in the set of documents are presented. ​​​​​​​​​​​It becomes possible. The user does not need to select keywords and can easily search for desired information from the document. and retrieve it.

[0067] In the document data processing system and the document data processing method of the present embodiment, a language model is used in which even the same word can obtain different distributed representations of words depending on the sentences containing it. By this, it is possible to search for a block highly relevant to the information required by the user with high accuracy.

[0068] The present embodiment can be appropriately combined with other embodiments. Also, in this specification, when a plurality of configuration examples are shown in one embodiment, it is possible to appropriately combine the configuration examples.

[0069] (Embodiment 2) In the present embodiment, a document data processing system according to an aspect of the present invention will be described with reference to FIGS. 6 and 7.

[0070] The document data processing system of the present embodiment can easily search for and obtain desired information from a document by using the document data processing method shown in Embodiment 1.

[0071] <Configuration Example 1 of Document Data Processing System> FIG. 6 shows a block diagram of a document data processing system 200. Note that in the drawings attached to this specification, the components are classified by function and the block diagram is shown with each component as an independent block. However, in actuality, it is difficult to completely separate the components by function, and one component may be related to multiple functions. Also, one function may be related to multiple components. For example, the processing performed by the processing unit 120 may be executed on different servers depending on the processing. ​​​​​​​It may be done.

[0072] The document data processing system 200 has at least a processing unit 120. The document shown in FIG. 6 The data processing system 200 further has an input unit 110, a storage unit 130, a database 14 0, a display unit 150, and a transmission path 160.

[0073] [Input Unit 110] A query sentence (query text) is supplied to the input unit 110 from outside the document data processing system 200. Also, a set of target documents may be supplied to the input unit 110 from outside the document data processing system 200. The set of target documents and the query sentence supplied to the input unit 110 are each supplied to the processing unit 120, the storage unit 130, or the database 140 via the transmission path 160.

[0074] The target document and the query sentence are input, for example, as text data, voice data, or image data. The target document is preferably input as text data.

[0075] Examples of the input method of the query sentence include key input using a keyboard, touch panel, etc., voice input using a microphone, reading from a recording medium, image input using a scanner, camera, etc., and acquisition using communication.

[0076] The document data processing system 200 may have a function of converting voice data into text data. For example, the processing unit 120 may have such a function. Or, the document data processing system 200 may further have a voice conversion unit having such a function.

[0077] ​​The document data processing system 200 may have an optical character recognition (OCR) function. Thus, it is possible to recognize characters included in the image data and create text data. For example, the processing unit 120 may have the said function. Or, the document data processing system 200 may further have a character recognition unit having the said function.

[0078] [Processing Unit 120] The processing unit 120 has a function of performing calculations using data supplied from the input unit 110, the storage unit 130, the database 140, etc. The processing unit 120 can supply the calculation result to the storage unit 130, the database 140, the display unit 150, etc.

[0079] The processing unit 120 has a function of dividing a document into a plurality of blocks. For example, it may have a function of dividing a document into a plurality of blocks such as by chapter, by paragraph, or by a predetermined number of sentences.

[0080] The processing unit 120 has a function of obtaining a distributed representation of words. For example, it is possible to obtain the distributed representation of words included in the blocks of the target document or words included in the query sentence.

[0081] The processing unit 120 has a function of extracting words from the query sentence. Thereby, among the words included in the query sentence, it is possible to select the words to be used for calculating the similarity.

[0082] The processing unit 120 has a function of calculating the similarity between the distributed representations of words.

[0083] For the processing unit 120, a transistor having a metal oxide in the channel formation region may be used. Since the off-current of the said transistor is extremely low, using the said transistor as a memory element By using it as a switch for holding the charge (data) flowing into the capable capacitive element , the data holding period can be ensured over a long period. By applying this characteristic to at least one of the register and the cache memory that the processing unit 120 has , the processing unit 120 is operated only when necessary, and in other cases, the information of the previous process is saved in the storage element , so that the processing unit 120 can be turned off. That is, normal-off operation is enabled, and power consumption of the document data processing system can be reduced. .

[0084] In this specification and the like, a transistor using an oxide semiconductor in the channel formation region is called an O xide Semiconductor transistor, or an OS transistor. The channel formation region of the OS transistor preferably has a metal oxide.

[0085] The metal oxide included in the channel formation region preferably contains indium (In). When the metal oxide included in the channel formation region is a metal oxide containing indium , the carrier mobility (electron mobility) of the OS transistor increases. Further, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing an element M. The element M is preferably aluminum (Al) , gallium (Ga), or tin (Sn). Other elements applicable to the element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni) , germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf) . There are hafnium (Hf), tantalum (Ta), tungsten (W), etc. However, as the element M , there may be cases where a plurality of the aforementioned elements can be combined. The element M is, for example, an element with a high binding energy with oxygen. For example, it is an element with a higher binding energy with oxygen than indium. Also, the metal oxide included in the channel formation region preferably contains zinc (Zn). The metal oxide containing zinc may be likely to crystallize. The metal oxide included in the channel formation region is not limited to the metal oxide containing indium. The semiconductor layer may be, for example, a metal oxide that does not contain indium, such as zinc tin oxide or gallium tin oxide, and contains zinc, a metal oxide containing gallium, a metal oxide containing tin, etc.

[0086] The metal oxide included in the channel formation region is not limited to the metal oxide containing indium. The semiconductor layer may be, for example, a metal oxide that does not contain indium, such as zinc tin oxide or gallium tin oxide, and contains zinc, a metal oxide containing gallium, a metal oxide containing tin, etc. The metal oxide included in the channel formation region is not limited to the metal oxide containing indium. The semiconductor layer may be, for example, a metal oxide that does not contain indium, such as zinc tin oxide or gallium tin oxide, and contains zinc, a metal oxide containing gallium, a metal oxide containing tin, etc.

[0087] Also, for the processing unit 120, a transistor including silicon in the channel formation region may be used. Also, for the processing unit 120, a transistor including silicon in the channel formation region may be used.

[0088] Also, for the processing unit 120, a transistor including an oxide semiconductor in the channel formation region and a transistor including silicon in the channel formation region may be used in combination. Also, for the processing unit 120, a transistor including an oxide semiconductor in the channel formation region and a transistor including silicon in the channel formation region may be used in combination.

[0089] The processing unit 120 has, for example, an arithmetic circuit or a central processing unit (CPU: Central Processing Unit), etc. The processing unit 120 has, for example, an arithmetic circuit or a central processing unit (CPU: Central Processing Unit), etc.

[0090] The processing unit 120 may have a microprocessor such as a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The processing unit 120 may have a microprocessor such as a DSP (Digital Signal Processor) or a GPU (Graphics Processing Unit). The microprocessor may be an FPGA (Field Programmable a Gate Array), FPAA (Field Programmable A nalog Array), etc., may be realized by a PLD (Programmable Logic Dev ice). The processing unit 120 interprets and executes instructions from various programs by a processor, and can perform various data processing and program control . Programs executable by the processor are stored in at least one of the memory area of the processor and the storage unit 130 . The processing unit 120 may have a main memory. The main memory has at least one of a volatile memory such as a RAM and a non-volatile memory such as a ROM

[0091] . As the RAM, for example, DRAM (Dynamic Random Access Me mory), SRAM (Static Random Access Memory), etc. are used, and a memory space is virtually allocated and used as the working space of the processing unit 120

[0092] . The operating system, application programs , program modules, program data, and look-up tables stored in the storage unit 130 are loaded into the RAM for execution . These data, programs, and program modules loaded into the RAM are directly accessed and operated on by the processing unit 120 . . .

[0093] . The ROM can store a BIOS (Basic Input / Outpu t System) and firmware, etc., which do not require rewriting. As the ROM, a mask ROM, an OTPROM (One Time Programmable Read Only Memory, etc. can be used Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. As for EPROM, there are UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) which enables erasure of stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmabl e Read Only Memory), flash memory, etc.

[0094] [Memory unit 130] The memory unit 130 has a function of storing the program executed by the processing unit 120. Also, the memory unit 130 may have a function of storing, for example, the calculation result generated by the processing unit 120 and the data input to the input unit 110. Specifically, the memory unit 130 preferably has a function of storing the distributed representation of words acquired by the processing unit 120.

[0095] The memory unit 130 has at least one of a volatile memory and a non-volatile memory. The memory unit 130 may have, for example, a volatile memory such as DRAM or SRAM. The memory unit 130 may have, for example, ReRAM (Resistive Random Access Memory, also called resistive change memory), PRAM (Phase change R andom Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresis tive Random Access Memory, also called magnetic resistance memory), ​​​It may also have a non-volatile memory such as a flash memory. Also, the storage unit 130 may have a recording medium drive such as a Hard Disc Drive (HDD) and a Solid State Drive (SSD).

[0096] [Database 140] The document data processing system may have a database 140. For example, the database 140 may have a function of storing a plurality of documents. For example, the document data processing method according to one aspect of the present invention may be used for a set of documents stored in the database 140. Note that the storage unit 130 and the database 140 may not be separated from each other. For example , the document data processing system may have a storage unit having the functions of both the storage unit 130 and the database 140.

[0097] Note that the memories of the processing unit 120, the storage unit 130, and the database 140 can each be regarded as an example of a non-transitory computer-readable storage medium.

[0098] [Display unit 150] The display unit 150 has a function of displaying the calculation result in the processing unit 120. Also, the display unit 150 has a function of displaying a target document. Also, the display unit 150 may have a function of displaying a query sentence.

[0099] Note that the document data processing system 200 may have an output unit. The output unit has a function of supplying data to the outside.

[0100] [Transmission path 160] ​​​​​​The transmission path 160 has a function of transmitting various data. The data transmission and reception between the input unit 110, the processing unit 120, the memory unit 130, the database 140, and the display unit 150 can be performed via the transmission path 1 60. For example, data such as a target document is transmitted and received via the transmission path 160.

[0101] <Configuration Example 2 of Document Data Processing System> FIG. 7 shows a block diagram of a document data processing system 210. The document data processing system 2 10 includes a server 220 and a terminal 230 (such as a personal computer).

[0102] The server 220 includes a communication unit 161a, a transmission path 162, a processing unit 120, and a memory unit 170. Although not shown in FIG. 7, the server 220 may further include an input / output unit and the like.

[0103] The terminal 230 includes a communication unit 161b, a transmission path 164, a processing unit 180, a memory unit 130, and a display unit 150. Although not shown in FIG. 7, the terminal 230 may further include a database and the like.

[0104] The user of the document data processing system 210 inputs a question sentence (query sentence) to the input unit 110 of the terminal 230. The question sentence is transmitted from the communication unit 161b of the terminal 230 to the communication unit 161a of the server 220.

[0105] The question sentence received by the communication unit 161a is stored in the memory unit 170 via the transmission path 162. Alternatively, the question sentence may be directly supplied from the communication unit 161a to the processing unit 120.

[0106] The document splitting, distributed representation acquisition, and similarity calculation described in Embodiment 1 are each highly​​​ The processing unit 120 of the server 220 is required to have the same processing power as the terminal 230. Therefore, these processes are performed by the processing unit 12. It is preferably done at 0.

[0107] Then, the processor 120 calculates the score of the block. Alternatively, the score may be directly transmitted from the processing unit 120 to the storage unit 170 via the The score may be supplied from the communication unit 161a of the server 220 to the terminal 23. The score is transmitted to the communication unit 161b of the terminal 230. The score is displayed on the display unit 150 of the terminal 230.

[0108] [Transmission path 162 and transmission path 164] The transmission path 162 and the transmission path 164 have a function of transmitting data. Data transmission and reception between the control unit 120 and the storage unit 170 is performed via a transmission path 162. The input unit 110, the communication unit 161b, the processing unit 180, the storage unit 130, and the display unit 1 Data transmission and reception between the devices 50 can be performed via a transmission path 164.

[0109] [Processing Unit 120 and Processing Unit 180] The processing unit 120 uses data supplied from the communication unit 161a and the storage unit 170, etc. The processing unit 180 has a function of performing calculations. The processing unit 120 and the processing unit 150 perform calculations using data supplied from the processing unit 120 and the processing unit 150. The processing unit 180 can refer to the description of the processing unit 120. It is preferable that all of them have high processing capacity.

[0110] [Storage section 130] The storage unit 130 has a function of storing programs executed by the processing unit 180. Also, the storage unit 130 has a function of storing calculation results generated by the processing unit 180, data input to the communication unit 161b, and data input to the input unit 110, etc.

[0111] [Storage unit 170] The storage unit 170 has a function of storing a plurality of documents, calculation results generated by the processing unit 120, and data input to the communication unit 161a, etc.

[0112] [Communication unit 161a and communication unit 161b] Using the communication unit 161a and the communication unit 161b, data can be transmitted and received between the server 220 and the terminal 230. As the communication unit 161a and the communication unit 161b, a hub, a router, a modem, etc. can be used. For data transmission and reception, either wired or wireless (e.g., radio waves, infrared rays, etc.) can be used.

[0113] Note that the communication between the server 220 and the terminal 230 can be performed by connecting to a computer network such as the Internet, which is the basis of the World Wide Web (WWW), an intranet, an extranet, a PAN (Personal Area Network), a LAN (Local Area Network ), a CAN (Campus Area Network), a MAN (Metropolitan Area Network ), a WAN (Wide Area Network ), a GAN (Global Area Network), etc.

[0114] This embodiment can be appropriately combined with other embodiments.

Explanation of reference numerals

[0115] W1: Word, W2: Word, 1: Block, 2: Block, 3: Block, 4: Block, 100: Document data processing system, 101: Document reading unit, 102: Question sentence input unit, 103 : Document splitting unit, 104a: Distributed representation acquisition unit, 104b: Distributed representation acquisition unit, 105a: Distributed representation holding unit, 105b: Distributed representation holding unit, 106: Word selection unit, 107: Similarity calculation unit, 108: Score display unit, 109: Sentence display unit, 110: Input unit, 120: Processing unit, 130 : Memory unit, 140: Database, 150: Display unit, 160: Transmission path, 161a: Communication unit , 161b: Communication unit, 162: Transmission path, 164: Transmission path, 170: Memory unit, 180: Processing unit, 200: Document data processing system, 210: Document data processing system, 220: Server , 230: Terminal

Claims

[Claim 1] a document reading unit that reads a plurality of target documents; a document division unit that divides each of the plurality of target documents into a plurality of blocks; a first distributed representation acquisition unit that acquires distributed representations of words for each block; a first distributed representation holding unit that stores the distributed representations acquired by the first distributed representation acquisition unit for each of the target documents and for each of the blocks; a query sentence reading unit that reads a query sentence; a second distributed representation acquisition unit that extracts words included in the query sentence and acquires distributed representations of the words included in the query sentence; a second distributed representation storage unit that stores the distributed representations acquired by the second distributed representation acquisition unit; and a similarity calculation unit that compares the embedded representations of words included in the query sentence with the embedded representations of words included in each of the plurality of blocks and calculates the similarity for each of the blocks; A document data processing system including: The similarity calculation unit searches for words included in the block that match words included in the query sentence, and calculates the similarity between the embedded representation of the word in the block and the embedded representation of the word in the query sentence for the matching words.

Citation Information

Patent Citations

  • Retrieval device, similarity calculation method, and program

    JP2019082931A

  • Document search using grammatical units

    US20190155913A1

  • Document reading comprehension support device, document reading comprehension support system, and program

    JP2014219833A

  • JP2021550716A