Document search system

The document search system addresses the limitation of conventional word-based searches by creating conceptual graph structures to enhance search accuracy and relevance through advanced analysis techniques, enabling more precise document retrieval.

JP2026076330APending Publication Date: 2026-05-11SEMICON ENERGY LAB CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SEMICON ENERGY LAB CO LTD
Filing Date
2026-02-17
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Conventional document search methods primarily rely on single-word searches, failing to account for the conceptual similarity between documents, especially in fields like patents and contracts where similar words abound, necessitating a more accurate search technology that considers document concepts.

Method used

A document search system that analyzes documents to create graph structures, incorporating part-of-speech tagging, dependency analysis, and token abstraction to generate conceptual graphs, using techniques like Weisfeiler-Lehman kernels for similarity evaluation.

Benefits of technology

Enhances search accuracy by considering document concepts, improving the ranking and relevance of search results by analyzing sentence structures and relationships, reducing reliance on word-level similarities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026076330000001_ABST
    Figure 2026076330000001_ABST
Patent Text Reader

Abstract

Search for documents, taking the concept of a document into consideration. [Solution] The device comprises an input unit, a first processing unit, a storage unit, a second processing unit, and an output unit. The input unit has the function of inputting the first document, and the first processing unit takes the first document and processes it. The first part has the function of creating a graph structure, and the second part has the function of storing the second graph structure. The second processing unit has a function to calculate the similarity between the first graph structure and the second graph structure. The output unit has the function of supplying information, and the first processing unit outputs the first document to multiple It has the function of dividing into tokens, and the nodes and edges of the first graph structure have labels. Furthermore, a label is composed of multiple tokens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One aspect of the present invention relates to a document search system. Another aspect of the present invention relates to a method for searching for documents.

Background Art

[0002] Various search techniques for searching documents are provided. In conventional document searches, single-word (string) searches are mainly used. For example, page rank and the like are used in web pages, and thesaurus is used in the patent field. Also, there is a method of expressing the similarity of documents by using sets of words and using Jaccard coefficient, Dice coefficient, Simpson coefficient, etc. Further, there are techniques such as vectorizing documents using tf-idf, Bag of Words (BoW), Doc2Vec, etc. and comparing cosine similarities.

[0003]

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] ​To search for documents across various fields, a more accurate document search method is required. In documents such as patent documents (specifications, claims, etc.) and contracts, there are many similar words. It is often used. Therefore, it is important to consider not only the words used in the document, but also the concept of the document. Search technology will be crucial.

[0005] Therefore, one aspect of the present invention provides a document search system that takes into account the concept of a document. This is one of the challenges. Furthermore, one aspect of the present invention is a method for searching for documents, taking into account the concept of a document. One of our objectives is to provide legal frameworks.

[0006] Furthermore, the description of these problems does not preclude the existence of other problems. One approach does not require that all of these issues be resolved. The title will become clear from the description in the specification, drawings, claims, etc. It is possible to extract other issues from the descriptions in the drawings, claims, etc. [Means for solving the problem]

[0007] One aspect of the present invention comprises an input unit, a first processing unit, a storage unit, a second processing unit, and an output unit. It is a document search system having the following: The input unit has the function of inputting a first document, and the first The processing unit has the function of creating a first graph structure from a first document, and the storage unit has the function of creating a second graph structure from a first document. The second processing unit has the function of storing the graph structure, and the first graph structure and the second graph The first has a function to calculate the similarity between the structure and the output unit, and has a function to supply information. The processing unit has the function of dividing the first document into multiple tokens, and the first graph structure Nodes and edges have labels, and labels consist of multiple tokens.

[0008] In the above document search system, it is preferable that the first processing unit has a function of assigning a part-of-speech to a token. Doing so is preferable.

[0009] Also, in the above document search system, it is preferable that the first processing unit has a function of performing dependency analysis. And the first processing unit has a function of concatenating a part of tokens according to the result of the dependency analysis. Doing so is preferable.

[0010] Also, in the above document search system, it is preferable that the first processing unit has a function of replacing a token having a representative word or a hypernym with the representative word or the hypernym. Doing so is preferable.

[0011] Also, in the above document search system, it is preferable that the second graph structure is created from the second document by the first processing unit. Doing so is preferable.

[0012] Also, in the above document search system, when the label of an edge of the graph structure has an antonym, the first processing unit has a function of generating a new graph structure by reversing the direction of the edge of the graph structure and replacing the label of the edge with the antonym. Doing so is preferable. Doing so is preferable. Doing so is preferable.

[0013] Also, in the above document search system, it is preferable that the second processing unit has a function of vectorizing the first graph structure and the second graph structure and evaluating the vector similarity between the vectorized first graph structure and the vectorized second graph structure. Doing so is preferable. Doing so is preferable. Doing so is preferable.

[0014] Also, in the above document search system, it is preferable that the second processing unit has a function of vectorizing the first graph structure and the second graph structure and evaluating the vector similarity between the vectorized first graph structure and the vectorized second graph structure, and... Vectorize the graph structure of 2 using the Weisfeiler-Lehman kernel It preferably has the function.

[0015] Also, in the above document search system, the part-of-speech assigned to the first token is a noun and the part-of-speech assigned to the second token located immediately before the first token is an adjective In this case, it is preferable that the first processing unit has a function of concatenating the second token and the first token It is preferable.

[0016] Also, in the above document search system, when both the part-of-speech assigned to the third token and the part-of-speech assigned to the fourth token located immediately after the third token are nouns it is preferable that the first processing unit has a function of concatenating the third token and the fourth token It is preferable. It is preferable.

Advantages of the Invention

[0017] According to one aspect of the present invention, it is possible to provide a document search system that takes into account the concepts of documents Also, according to one aspect of the present invention, it is possible to provide a method for searching documents that takes into account the concepts of documents It is possible.

[0018] By analyzing each sentence of a document to obtain a conceptual graph structure and calculating the similarity of the graph structures, it is possible to search for documents that are conceptually close In addition, by combining with conventional search methods, the accuracy such as ranking can be improved It can be improved.

[0019] Note that the effects of one aspect of the present invention are not limited to the effects listed above. The effects listed above do not prevent the existence of other effects. Note that other effects will be described in the following description, the present This is an effect not mentioned in the section. An effect not mentioned in this section would be obvious to someone skilled in the art. This can be derived from detailed descriptions, drawings, etc., and can be extracted appropriately from these descriptions. Yes, it is possible. Furthermore, one aspect of the present invention includes, among the effects listed above and / or other effects, at least one of the following: It has at least one effect. Therefore, one aspect of the present invention may, in some cases, It may not always have the effects listed above. [Brief explanation of the drawing]

[0020] [Figure 1] Figure 1 shows an example of a document search system. [Figure 2] Figure 2 is a flowchart showing an example of how to search for documents. [Figure 3] Figures 3A to 3C show the results obtained in each step. [Figure 4] Figures 4A to 4C show the results obtained in each step. [Figure 5] Figures 5A to 5D show the results obtained in each step. [Figure 6] Figures 6A to 6C show the results obtained in each step. [Figure 7] Figure 7 shows an example of the hardware for a document search system. [Figure 8] Figure 8 shows an example of the hardware for a document search system. [Modes for carrying out the invention]

[0021] Embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description. Without departing from the spirit and scope of the present invention, its form and details may be changed in various ways. Those skilled in the art will readily understand that further improvements are possible. Therefore, the present invention can be implemented as follows: It is not to be interpreted as being limited to the description of the form.

[0022] Furthermore, in the configuration of the invention described below, the same part or part having a similar function The same reference numeral is used in common across different drawings, and explanations of its repetition are omitted. When referring to a specific function, the hatch pattern may be the same, and no special designation may be assigned.

[0023] Furthermore, the position, size, and scope of each component shown in the drawings are, for the sake of ease of understanding, actual The position, size, and range of the edges may not be shown. Therefore, the disclosed invention is not necessarily However, this is not limited to the location, size, and scope disclosed in the drawings.

[0024] Furthermore, the ordinal numbers "1st," "2nd," and "3rd" used in this specification refer to the constituent elements. This note is added to avoid confusion and does not imply any numerical limitation.

[0025] (Embodiment 1) This embodiment provides a document search system and a method for searching for documents according to one aspect of the present invention. This will be explained using Figures 1 through 4C.

[0026] <Document Search System> Figure 1 is a diagram showing the configuration of the document search system 100. In other words, Figure 1 shows one of the present inventions This can be considered an example of the configuration of a document search system.

[0027] The document search system 100 is an information processing system that uses personal computers and other devices used by users. It may be provided in the device. Alternatively, the processing unit of the document search system 100 may be provided in the server. Alternatively, it may be configured so that it is accessed and used from a client PC via the network.

[0028] As shown in Figure 1, the document search system 100 includes an input unit 101 and a graph structure creation unit 10 2. It comprises a similarity calculation unit 103, an output unit 104, and a storage unit 105. The unit includes a graph structure creation unit 102 and a similarity calculation unit 103.

[0029] The input unit 101 inputs document 20. Document 20 is a document specified by the user for search purposes. Yes. Document 20 is text data, audio data, or image data. Input unit 10 1. Input devices such as keyboards, mice, touch sensors, microphones, scanners, and cameras. There is a vise.

[0030] The document search system 100 has a function to convert audio data into text data. Alternatively, the graph structure creation unit 102 may have this function. The search system 100 may also have a speech-to-text conversion unit having the said function. stomach.

[0031] The document search system 100 may also have an optical character recognition (OCR) function. This allows for the recognition of characters contained in image data and the creation of text data. For example, the graph structure creation unit 102 may have that function. Alternatively, the document search system The M100 may further have a character recognition unit having the said function.

[0032] The storage unit 105 stores documents 10_1 to 10_n (where n is an integer of 2 or more). Documents 10_1 through 10_n are comparison documents for document 20. In some cases, documents 10_1 through 10_n may be collectively referred to as multiple documents 10. Multiple documents 10 are stored in the storage unit 105 via the input unit 101, storage medium, communication, etc. It can be done.

[0033] The multiple documents 10 stored in the storage unit 105 are preferably text data. For example, by converting audio data or image data into text data, The size can be reduced, and the load on the storage unit 105 can be reduced.

[0034] Furthermore, the storage unit 105 stores graph structures 11_1 to 11_n. Graph structures 11_1 through 11_n correspond to documents 10_1 through 10_n, respectively. This is the graph structure in contrast. Note that graph structures 11_1 to 11_n are each Then, from documents 10_1 to 10_n, the graph structure is created by the graph structure creation unit 102. So, let's group graph structures 11_1 to 11_n together and consider multiple graph structures 11 and It may be written as such.

[0035] Document 10_i (where i is an integer between 1 and n, inclusive) and graph structure 11_i are identical. It is preferable that an ID is assigned to it. This allows document 10_i and graph structure 1 1_i can be associated with graph structures 11_1 to 11_n. By creating these in advance, you can reduce the time required to search for documents.

[0036] Furthermore, the storage unit 105 may store the document 20. Graph structure 21 may be stored. Note that graph structure 21 is created from document 20. It is created in Narube 102.

[0037] The graph structure creation unit 102 has the function of creating a graph structure from a document. The rough structure creation unit 102 has functions for morphological analysis, dependency parsing, and abstraction. It is preferable that it has the function and the function to create a graph structure. Section 102 has the function of referencing the conceptual dictionary 112. Referencing the conceptual dictionary 112, the graph In the structure creation unit 102, a graph structure is created for the document. The document in question is document 20. and multiple documents 10.

[0038] The graph structure is preferably a directed graph. A directed graph is one in which nodes and directions are directed. It is a graph composed of nodes and edges. Furthermore, the graph structure consists of nodes and edges. It is more preferable that the graph is a directed graph with labels assigned to it. By using the graph structure of graphs, similarity and search accuracy can be improved. .

[0039] Note that in Figure 1, the conceptual dictionary 112 is installed in a device different from the document search system 100. The configuration shown is, but is not limited to, this. Conceptual dictionary 112 is a document search system It may be prepared for 100.

[0040] Furthermore, the functions for morphological analysis and dependency parsing are provided by Document Search System 1 It may be provided in a device different from 00. In this case, the document search system 100 is the above sentence The document is transmitted to the device, and the results of the morphological analysis and dependency parsing performed by the device are then processed. The system receives data and transmits the received data to the graph structure creation unit 102.

[0041] The similarity calculation unit 103 calculates the similarity between the first graph structure and the second graph structure. It has a function. The first graph structure is graph structure 21. The second graph structure is multiple It is one or more of the graph structures 11. In other words, the similarity calculation unit 103 calculates the first The similarity between the first document and the second document is evaluated. The first document is document 20. The second The document is one or more of several documents 10.

[0042] The output unit 104 has the function of supplying information. This information is calculated by the similarity calculation unit 103. This is information regarding the calculated similarity results. For example, this information is used for multiple documents 10. This document has the highest similarity to document 20. Alternatively, this information is similar to document 10_i, The similarity between document 20 and document 10_i, and the pairs, sorted in descending order of similarity. In this case, the number of pairs is between 2 and n.

[0043] The above information is supplied as, for example, text strings, numbers, visual information such as graphs, and audio information. Output devices such as a display and speakers are included as output section 104.

[0044] The document search system 100 has a function to convert text data into audio data. It is also possible that the document search system 100 further has the function of text-to-speech conversion It may have a replacement part.

[0045] The above is a description of the configuration of the document search system 100. This is one aspect of the present invention. By using the document search system, documents conceptually similar to document 20 can be found among multiple documents 10. You can search within that list. Additionally, you can view a list of documents conceptually similar to document 20, and multiple similar documents. You can search within document 10.

[0046] According to one aspect of the present invention, a document search system that takes into account the concept of a document can be provided. Cut.

[0047] <How to search for documents> Figure 2 is a flowchart illustrating the processing flow performed by the document search system 100. In other words, Figure 2 shows a flowchart illustrating an example of a method for searching for documents according to one aspect of the present invention. It can also be said to be a techno.

[0048] In one embodiment of the present invention, a method for searching documents involves analyzing the documents and converting them into a graph structure, The similarity of graph structures is measured using the Weisfeiler-Lehman (WL) kernel, etc. By comparing them, you can search for documents.

[0049] Step S001 is the process of obtaining multiple documents 10. The multiple documents 10 are stored These are documents stored in section 105. Multiple documents 10 are stored in the input section 101, storage medium, and It is stored in storage unit 105 via a signal or other means.

[0050] If multiple documents 10 constitute patent claims, before proceeding to step S002 Alternatively, document cleaning may be performed on each of the multiple documents 10. Leaning involves actions such as removing semicolons or replacing colons with commas. Yes, it is possible. Document cleaning can improve the accuracy of morphological analysis.

[0051] Furthermore, the cleaning of the above documents is performed when multiple documents 10 are outside the scope of the patent claims. Even in this case, it is best to do so as needed. Also, multiple documents 10 are the same as the above document. After cleaning, it may be stored in the storage unit 105.

[0052] Step S002 is performed by the graph structure creation unit 102 for each of the multiple documents 10. This is the process of performing morphological analysis. As a result, each of the multiple documents 10 is divided into morphemes. They are divided. In this specification, the divided morphemes may be referred to as tokens.

[0053] In step S002, for each of the above divided morphemes (tokens), It is preferable to identify the part of speech of each morpheme (token) and associate it with a part-of-speech label. By associating part-of-speech labels with tokens, the accuracy of dependency parsing can be improved. This is possible. Furthermore, in this specification, etc., we associate morphemes (tokens) with part-of-speech labels. This can be rephrased as assigning a part of speech to a morpheme (token).

[0054] If the graph structure creation unit 102 does not have a function to perform morphological analysis, the document search system Using a morphological analysis program (also called a morphological analyzer) incorporated into a different device Alternatively, morphological analysis may be performed on each of the multiple documents 10. In this case, step S002 transmits multiple documents 10 to the device, and the device performs morphological analysis, and the morphology This is the process of receiving the data resulting from the elementary analysis.

[0055] Step S003 is the process of performing dependency analysis in the graph structure creation unit 102. In other words, depending on the dependency of each divided morpheme (token), multiple tokens This is the process of combining parts of something. For example, if a token satisfies certain conditions, the conditions Combine existing tokens to generate a new token.

[0056] If Japanese is used in the document, specifically, the jth (where j is an integer greater than or equal to 2). The token is a noun, and the token immediately preceding the jth token is (the (j-1) This is called a token. If the adjective is a (j-1) token and the j-th token. Combine the two to generate a new token. Also, the jth token is a noun and The token immediately following the jth token (called the (j+1)th token) is a noun. If so, combine the j-th token and the (j+1)th token to form a new token Generates.

[0057] The above conditions should be adjusted as needed to match the language used in the document.

[0058] The above dependency parsing preferably includes compound word parsing. By doing so, parts of multiple tokens are combined to create a new token, generating a compound word. This allows for the creation of compound words in a document that are not registered in the conceptual dictionary 112. However, the document can be divided into tokens with high accuracy.

[0059] If the graph structure creation unit 102 does not have a function to perform dependency parsing, the document search system A dependency parsing program (also called a dependency parser) incorporated into a device different from the one described. Dependency parsing may be performed using this method. In this case, step S003 is divided into The token is transmitted to the device, and the device performs dependency parsing. This is the process of receiving the resulting data.

[0060] Step S004 is the process of abstracting tokens in the graph structure creation unit 102. For example, analyze the words contained in the token to obtain a representative word. Also, for that representative word... If a superordinate term exists, acquire that superordinate term. Then, use that token as the acquired representative. Replace with the word or its superordinate term. Here, the representative word is the headword of the group of synonyms. (Also called a lemma.) Furthermore, a hypernomial term is a representative term that corresponds to a higher-level concept of a representative term. Yes, it exists. In other words, token abstraction means replacing tokens with representative words or superordinate words. This refers to [something]. Note that if a token is a representative word or a hypernomial word, that token will not be replaced. That's fine.

[0061] The upper limit of the hierarchy of superordinate words to be replaced is preferably between 1 and 2, and is 1. This is preferable. Furthermore, the upper limit of the hierarchy of superordinate terms to be replaced may be specified. This helps to prevent tokens from being overly conceptualized.

[0062] The appropriate level of abstraction for tokens varies depending on the field. Therefore, machine learning tailored to each field... It is preferable to abstract the token by doing so. Token abstraction is, for example, The process involves vectorizing the token using its morphemes and then classifying it using a classifier. The classifiers used will include decision trees, support vector machines, and random number generators. Algorithms such as Forest and multilayer perceptrons may be used. Specifically, "oxidation "Solid semiconductor," "amorphous semiconductor," "silicon semiconductor," and "GaAs semiconductor" It would be best to classify these as "semiconductors." Also, "oxide semiconductor layers" and "oxide semiconductor films" "Amorphous semiconductor layer", "Amorphous semiconductor film", "Silicon semiconductor layer", "Silicon "Recon semiconductor film," "GaAs semiconductor layer," and "GaAs semiconductor film" are also classified as "semiconductors." It would be good to be similar.

[0063] Furthermore, a classifier is used to determine whether or not to extract morphemes contained in the token. This is also fine. For example, when abstracting the token "oxide semiconductor layer", the token can be expressed as It is broken down again into morphemes, and the decomposed morphemes are "oxidation", "matter", "semiconductor", and " Input "layer" into the classifier. If the result of inputting into the classifier is classified as "semiconductor", then Replace "-kun" with "semiconductor". This allows us to abstract the token. .

[0064] In addition to the above machine learning algorithms, there are also conditional random field algorithms. You may also use a (ndom field:CRF). Alternatively, you can combine CRF with the above method. It's okay to combine them.

[0065] By abstracting tokens, we can conceptually grasp documents. Therefore, sentences It is less affected by the structure and expression of the writing and allows for searching based on the conceptual factors of the document. .

[0066] Representative terms and hypernomies can be obtained using a conceptual dictionary or through machine learning-based classification. It is also acceptable to provide the concept dictionary in a device different from the document search system 100. You may use the concept dictionary 112 provided, or the concept dictionary provided in the document search system 100. You may use it.

[0067] Step S005 involves the graph structure creation unit 102 creating multiple graph structures 11. This is the process. In other words, the tokens prepared up to step S004 are used by the node or E In this context, it is the process of creating a graph structure. Specifically, in a document, the first is a noun phrase. This represents the relationship between the first token and the second token, and between the first token and the second token. If there is a third token, then the first token and the second token are each used by the node. The third token is used as the label of the node and the edge and the label of that edge. Create a graph structure. That is, the labels of the nodes and the labels of the edges are in step S. It consists of tokens prepared up to 004.

[0068] For example, if the document is a patent claim, the nodes in the graph structure are constituent elements It is a prime element, and the edges of the graph structure represent the relationships between its constituent elements. Also, documents such as contract documents... In some cases, the nodes of the graph structure are A and B, and the edges of the graph structure are determined by specific conditions. be.

[0069] The graph structure can be created based on rules derived from the dependency relationships between tokens. Additionally, CRF is used to assign labels to nodes and edges based on a list of tokens. You may perform machine learning to do this. This will allow you to use a list of tokens to determine the nodes and et Labels can be assigned to the data. Also, recurrent neural networks (Recur rent Neural Network:RNN), long short-term memory (Long sho Using rt-term memory (LSTM), etc., input a list of tokens. Alternatively, you can train a Seq2Seq model that outputs the orientation of nodes and edges. This allows us to output the orientation of nodes and edges from the list of tokens.

[0070] The graph structure creation unit 102 reverses the orientation of the edges and assigns labels to the edges. It may have a function to replace the label of the edge with its antonym. For example, if the graph structure is Edge 1 and the second edge which is assigned a label that is the antonym of the label of the first edge. If a edge and a label are present, reverse the orientation of the second edge and label the second edge Perform the process of replacing the label of the second edge with its antonym (i.e., the label of the first edge). Therefore, a new graph structure may be created. This will cover conceptually identical structures. Therefore, it is less affected by the structure and expression of the document, and the conceptual factors of the document. You can perform searches using this method.

[0071] Furthermore, the above process should be performed on the edge that appears less frequently in the document. In other words, If the frequency of occurrence of the second edge is lower than or the same as the frequency of occurrence of the first edge, Reverse the orientation of the second edge, and change the label of the second edge to the label of the second edge It is a good idea to replace it with its antonym (i.e., the label of the first edge). This will allow you to This can reduce the frequency of creating new graph structures.

[0072] The order of steps S004 and S005 may be reversed. If the order of steps 4 and S005 is reversed, after the graph structure is created, the The nodes and edges included in the rough structure are abstracted. Therefore, step S004 and Even if the order of steps S005 is changed, the abstracted graph structure can still be created from the document. It is possible.

[0073] Steps S001 to S005 allow multiple graph structures to be generated from multiple documents 10. The structure 11 can be created. Steps S001 to S005 are similar It is preferable to perform this before calculating the degree. Multiple graph structures 11 are created in advance. By doing so, the time required to search for documents can be reduced.

[0074] Step S011 is the process of obtaining document 20. Document 20 is obtained from input unit 101 This is the entered document. Note that document 20 is a text file of audio or image data. If the data is not text data, before proceeding to step S012, convert document 20 to text data. Convert to text data. The audio data held by the graph structure creation unit 102 is converted to text data. Function to convert to text data, or speech-to-text conversion unit, or graph structure creation. It is preferable to use the optical character recognition (OCR) function or character recognition unit of unit 102.

[0075] If document 20 is a claim, before proceeding to step S012, You may perform the document cleaning described above on document 20. This improves the accuracy of morphological analysis. Note that the cleaning of the document is performed as follows: Even if document 20 is not within the scope of the patent claims, it will be done as appropriate as necessary. good.

[0076] Step S012 involves the graph structure creation unit 102 performing morphological analysis on the document 20. This is the process. Note that step S012 is the same process as step S002, You can refer to the explanation in step S002.

[0077] Step S013 is the process of performing dependency analysis in the graph structure creation unit 102. Note that step S013 is the same process as step S003, therefore step S00 You can refer to explanation 3.

[0078] Step S014 is the process of abstracting tokens in the graph structure creation unit 102. Note that step S014 is the same process as step S004, therefore step S0 You can refer to the explanation in 04.

[0079] Step S015 is the process of creating the graph structure 21 in the graph structure creation unit 102. Yes. Note that step S015 is the same process as step S005, so step You can refer to the explanation in S005.

[0080] Step S016 is performed by the similarity calculation unit 103, which calculates the similarity between document 20 and the multiple documents 10. This is the process of evaluating the similarity to the graph structure 21 and multiple graphs. Specifically, the graph structure 21 and multiple graphs Structure 11 is vectorized by the WL kernel, and the vectorized graph structure 21 and the vector The similarity between each of the multiple torified graph structures 11 and the vector is evaluated.

[0081] Step S017 is the process of outputting information using the output unit 104. This information is: This information pertains to the similarity result calculated by the similarity calculation unit 103.

[0082] The above is a description of the method for searching for documents. Searching for documents according to one aspect of the present invention. By using this method, you can search for documents that are conceptually similar to the document you specify for the search. It can also search for documents that are conceptually similar to the document specified for the search, in a ranked order. It is possible to do so. Furthermore, it is less affected by the structure and expression of the document, and is not influenced by the conceptual factors of the document. You can perform a search.

[0083] One aspect of the present invention provides a method for searching for documents that takes into account the concept of a document. can.

[0084] <<Example of creating a graph structure from a document>> Regarding the document search methods described above, here is an example of creating a graph structure from the documents. This will be explained using Figures 3A to 6C.

[0085] First, "The oxide semiconductor layer is located above the insulating layer (SANKABUTSUHAND OUTAISOU HA ZETSUENTAISOU NO JOUHOU NI A This explanation will use a document that uses Japanese, specifically "RU)" (see Figure 3A), as an example. The rounded rectangles shown in Figures 3B, 3C, and 4A are tokens, and below the rounded rectangles This section lists the part of speech assigned to the token.

[0086] First, by performing morphological analysis on the above document, the document is divided into tokens, and each token Assign a part of speech to the word (as shown in Figure 2, steps S002 and S012). As a result, the results shown in Figure 3B are obtained. Specifically, the above document states, “Oxidation (SA NKA)” (noun) | “BUTSU” (noun) | “Semiconductor (HANDOUTAI)” (Noun) | "SOU" (Noun) | "HA" (Particle) | "ZETSU" EN)" (noun) | "body (TAI)" (noun) | "layer (SOU)" (noun) | "of (NO) )”(particle)|“upper (JOUHOU)”(noun)|“in (NI)”(particle)|“aru ( It is divided into tokens, such as "ARU)" (verb), and each token is assigned a part of speech. It will be done.

[0087] Next, dependency parsing is performed (steps S003 and S013, as shown in Figure 2). As a result, the results shown in Figure 3C are obtained. Specifically, "oxidation (SANKA)" and “thing (BUTSU)”, “thing (BUTSU)”, and “semiconductor (HANDOUTA I) ``Semiconductor (HANDOUTAI)'' and ``Layer (SOU)'' are, The conditions described in S003 are met. Therefore, the four tokens ("oxidation (SANK A)”, “BUTSU”, “Semiconductor (HANDOUTAI)”, “Layer (SOU)” ) are joined together, and one token ("Oxide semiconductor layer (SANKABUTSUHANDOU It can be replaced with "TAISOU)" and also with "ZETSUEN" and "Body (TAI)", as well as "Body (TAI)" and "Layer (SOU)", are steps. The conditions described in S003 are met. Therefore, the three tokens ("Isolation (ZETSUE N)”, “body (TAI)”, “layer (SOU)”) are combined into one token ("insulator"). It can be replaced with "layer (ZETSUENTAISOU)"). This will result in the above sentence The text reads, “Oxide semiconductor layer (SANKABUTSUHANDOUTAISOU)” (noun) )|"HA" (particle)|"ZETSUENTAISOU" (noun)| "no" (particle) | "upper (JOUHOU)" (noun) | "ni" (particle) It becomes "aru" (a verb).

[0088] Next, we perform token abstraction (as shown in Figure 2, steps S004 and S01) 4). As a result, the results shown in Figure 4A are obtained. Specifically, the oxide semiconductor layer ( "Semiconductor (HANDOUTAI)" can be replaced with the superordinate term "semiconductor (HANDOUTAI)". Also, "insulator layer (ZETSUENTAISOU)" can be replaced with the superordinate term "insulator (ZETSUENTAI)". Also, "upper part (JOUHOU)" can be replaced with the representative term "above (UE)". As a result, the above document is abstracted as "“semiconductor (HANDOUTAI)” (noun) | “is (HA)” (particle) | “insulator (ZETSUENTAI)” (noun) | “of (NO)” (particle) | “above (UE) ” (noun) | “on (NI)” (particle) | “exists (ARU)” (verb)".

[0089] Next, create a graph structure (steps S005 and S015 shown in Figure 2 ). As a result, a result as shown in Figure 4B is obtained. Specifically, "semiconductor (HANDO UTAI)" and "insulator (ZETSUENTAI)" become the nodes of the graph structure and the labels of the corresponding nodes, and "above (UE)" becomes the edges of the graph structure and the labels of the corresponding edges.

[0090] Here, the antonym of "above (UE)" is "below (SHITA)". Therefore, reverse the arrow of the graph structure shown in Figure 4B, and replace the edge of the graph structure shown in Figure 4B and the label of the corresponding edge, which is "above (UE)", with "below (SHITA)", and a new graph structure as shown in Figure 4C can be generated. As a result, it is possible to cover the conceptually same structure.

[0091] The arrows shown in Figures 4B and 4C are from the nodes that appear earlier in the document (in the case of the above document, "semiconductor (HAND OUTAI)") to the nodes that appear later (in the case of the above document, "insulator (​​​​ It is illustrated that the arrow is pointing towards "ZETSUENTAI)"." In other words, the starting point of the arrow is the first The node that appears in the diagram is designated as the node that appears later, and the endpoint of the arrow is designated as the node that appears later. This is not the only way to define morphology. For example, based on semantic relationships between words, such as spatial relationships, The direction of the arrow may be determined. Specifically, the starting point of the arrow may be the label "Insulator (ZETS) Let the node be labeled "UENTAI)" and the endpoint of the arrow be labeled "Semiconductor (HANDOUT Let the nodes be defined as "AI)" and the edges between these nodes and the labels of those edges be defined as "Above You may create a graph structure called (UE). This allows for an intuitive understanding of the graph structure. It is possible. However, the method for determining the direction of the arrow is a unified method in the document search method. It needs to be done.

[0092] Based on the above, an abstracted graph structure can be created from the document described above.

[0093] Next, "A semiconductor device comprising:a n oxide semiconductor layer over an insu As an example, let's look at a document that uses English, such as "later layer." (See Figure 5A). Let me explain. Note that the rounded rectangles shown in Figures 5C, 5D, and 6A are tokens. Note that here we show an example where tokens are not assigned parts of speech, but if parts of speech are assigned to tokens... That's fine.

[0094] First, we will clean the above document. Here, we will remove semicolons. As a result, the results shown in Figure 5B are obtained.

[0095] Next, morphological analysis is performed on the above document to divide it into tokens. Steps S002 and S012 are shown in Figure 2. As a result, as shown in Figure 5C... The following results can be obtained. Specifically, the above document is “A”|“semiconductor ”|“device”|“comprising”|“an”|“oxide”|“se microconductor”|“layer”|“over”|“an”|“insula It becomes "tor"|"layer"".

[0096] Next, dependency parsing is performed (steps S003 and S013, as shown in Figure 2). As a result, the results shown in Figure 5D are obtained. Specifically, three tokens ("A") The “semiconductor” and “device” are combined into one token ( It can be replaced with "A semiconductor device". , four tokens ("an", "oxide", "semiconductor", "l ayer) is combined into one token ("an oxide semiconduct It can be replaced with "tor layer"). Also, three tokens ("an", The “insulator” and “layer” are combined into one token ("an in This can be replaced with "sulator layer"). This means the above document can be replaced with: ““A semiconductor device”|“comprising”|“ an oxide semiconductor layer”|“over”|“an This becomes the "insulator layer".

[0097] Next, we perform token abstraction (as shown in Figure 2, steps S004 and S01) 4). As a result, the results shown in Figure 6A are obtained. Specifically, “A semico The term "inductor device" can be replaced with the broader term "device." Also, “an oxide semiconductor layer” is “a se It can be replaced with the superordinate term "miconductor". Also, "an insula The term "tor layer" can be replaced with the superordinate term "an insulator." Therefore, the above document is “device”|“comprising”|“as Abstraction as "emiconductor" | "over" | "an insulator" It will be done.

[0098] Next, create the graph structure (as shown in Figure 2, steps S005 and S015) ). As a result, the results shown in Figure 6B are obtained. Specifically, “deveice”, The graph shows that "semiconductor" and "insulator" are respectively... These become the nodes of the structure and the labels of those nodes, “comprising” and “o Each of the "ver" represents an edge in the graph structure and the label of that edge.

[0099] Here, the antonym of "over" is "under". Therefore, the graph shown in Figure 6B The structural arrows are reversed, and the edges of the graph structure shown in Figure 6B and the labels of those edges are reversed. By replacing the "over" with "under," we can obtain the graph structure shown in Figure 6C. It is also possible to generate new ones. This allows us to cover conceptually identical structures.

[0100] The arrows shown in Figures 6B and 6C indicate the first node to appear in the document (in the case of the above document, “se From "miconductor"), subsequent nodes (in the above document, "insul") appear (in the case of the above document, "insul") It is illustrated that it is pointing towards the ator). In other words, the starting point of the arrow is the no that appears first. This is designated as "do," and the endpoint of the arrow is designated as the node that will appear later. In this embodiment, this It's not limited. For example, the direction of the arrow can be determined based on the semantic relationships between words, such as their spatial relationships. It may be set as follows: Specifically, the starting point of the arrow is set to the node whose label is "insulator". Let D be the node whose label is "semiconductor", and the endpoint of the arrow be the node whose label is "semiconductor". Create a graph structure where the edges between these nodes and the labels of those edges are set to "over". This is acceptable. This allows for an intuitive understanding of the graph structure. However, arrows The method for determining the orientation of documents needs to be standardized across different document search methods.

[0101] Based on the above, an abstracted graph structure can be created from the document described above.

[0102] Furthermore, the process of creating a graph structure from a document is described for documents that use Japanese, and Although I used documents written in English as examples, the language of the documents is not limited to Japanese and English. It's not possible. Languages ​​such as Chinese, Korean, German, French, Russian, and Hindi are available. The same process can be followed to create a graph structure from the documents used. It is possible.

[0103] This embodiment can be appropriately combined with other embodiments. Furthermore, this specification In cases where multiple configuration examples are shown within a single embodiment, the configuration examples may be combined as appropriate. It is possible to combine them.

[0104] (Embodiment 2) In this embodiment, a document search system according to one aspect of the present invention is described using Figures 7 and 8. explain.

[0105] The document search system of this embodiment uses the document search method shown in Embodiment 1. This allows for easy searching of documents.

[0106] <Example of document search system configuration 1> Figure 7 shows a block diagram of the document search system 200. Note that the drawings attached to this specification are... So, let's classify the components by function and show them as independent blocks in a block diagram. However, it is difficult to completely separate the actual components by function, and one component is It may involve multiple functions. Also, one function may involve multiple components. For example, the processing performed in processing unit 202 may be executed on different servers depending on the process. It can happen.

[0107] The document search system 200 has at least a processing unit 202. (Figure 7 shows the document search) System 200 further includes an input unit 201, a storage unit 203, a database 204, and a display unit. It has a 205 and a transmission line 206.

[0108] [Input section 201] The input unit 201 receives documents from outside the document search system 200. This is a document specified by the user for search purposes, and corresponds to document 20 shown in Embodiment 1. Multiple documents may be supplied to the input unit 201 from outside the document search system 200. The aforementioned documents are documents that can be compared with the above documents, and are the multiple documents shown in Embodiment 1. This corresponds to document 10. The above multiple documents supplied to the input unit 201 and the above documents each Then, via the transmission line 206, to the processing unit 202, the storage unit 203, or the database 204 It will be supplied.

[0109] The above multiple documents and the above documents are, for example, text data, audio data, or image data. It is preferred that the above multiple documents be entered as text data. It's nice.

[0110] The above document can be entered using, for example, a keyboard or touch panel. Force, voice input using a microphone, reading from recording media, scanner, camera, etc. Examples include image input and data acquisition using communication.

[0111] The document search system 200 has a function to convert audio data into text data. Alternatively, the document search system may have the function. The M200 may further have a voice conversion unit having the said function.

[0112] The document search system 200 may also have an optical character recognition (OCR) function. This allows for the recognition of characters contained in image data and the creation of text data. For example, the processing unit 202 may have the function. Or, the document search system 200 may have Furthermore, it may also have a character recognition unit having the said function.

[0113] [Processing 202] The processing unit 202 is supplied with power from the input unit 201, storage unit 203, database 204, etc. It has the function of performing calculations using the data. The processing unit 202 stores the calculation results in the storage unit 20 3. It can be supplied to the database 204, display unit 205, etc.

[0114] The processing unit 202 consists of the graph structure creation unit 102 and the similarity calculation unit 1 shown in Embodiment 1. Includes 03. That is, the processing unit 202 has a function to perform morphological analysis and a function to perform dependency parsing. It has the ability to abstract, and the ability to create graph structures.

[0115] The processing unit 202 may use a transistor having a metal oxide in the channel formation region. i. Because the off-current of the transistor is extremely small, the transistor can be used as a memory element. It is used as a switch to hold the charge (data) that has flowed into a capacitive element that functions in this way. This ensures that data can be retained for a long period of time. This characteristic is utilized by the processing unit 20 By using it in at least one of the registers and cache memory of 2, the necessary The processing unit 202 is activated only in certain cases; otherwise, the information from the previous processing is saved to the memory element. By doing so, the processing unit 202 can be turned off. In other words, normally This enables computing power and reduces the power consumption of the document search system 200. It is possible.

[0116] Furthermore, in this specification, etc., a transistor using an oxide semiconductor in the channel formation region is defined as It is called an Oxide Semiconductor transistor (OS transistor). The channel formation region of the S-transistor preferably contains a metal oxide.

[0117] The metal oxide in the channel-forming region preferably contains indium (In). If the metal oxide in the channel-forming region is an indium-containing metal oxide, the OS transition The carrier mobility (electron mobility) of the star increases. Also, the metal present in the channel-forming region The oxide preferably contains element M. Element M is aluminum (Al), gallium ( It is preferable that the element M is Ga (Ga) or tin (Sn). Other elements applicable to element M include These are boron (B), titanium (Ti), iron (Fe), nickel (Ni), and germanium (G). e) Yttrium (Y), Zirconium (Zr), Molybdenum (Mo), Lanthanum (L) a) Cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), Examples include tungsten (W). However, as element M, multiple elements mentioned above can be combined. In some cases, this is acceptable. Element M is, for example, an element with a high bond energy with oxygen. For example, it is an element whose bonding energy with oxygen is higher than that of indium. Also, channel-shaped The metal oxide in the constituent region preferably contains zinc (Zn). Some substances may become more prone to crystallization.

[0118] The metal oxides present in the channel-forming region are not limited to indium-containing metal oxides. The metal oxides present in the channel-forming region are, for example, zinc-tin oxide and gallium-tin oxide. Metal oxides that do not contain indium but contain zinc, metal oxides that contain gallium, etc. It is also acceptable if it is a metal oxide containing ions.

[0119] Furthermore, even if a transistor containing silicon in the channel formation region is used in the processing unit 202, good.

[0120] Furthermore, the processing unit 202 includes a transistor containing an oxide semiconductor in the channel formation region, and A transistor containing silicon in the channel formation region may be used in combination with this.

[0121] The processing unit 202 is, for example, an arithmetic circuit or a central processing unit (CPU: Central Processing Unit). It has a processing unit, etc.

[0122] The processing unit 202 is a DSP (Digital Signal Processor), G Microprocessors such as PUs (Graphics Processing Units) It may have. Microprocessors are FPGAs (Field Programmers). ble Gate Array), FPAA (Field Programmable PLDs (Programmable Logic Decoders) such as Analog Arrays The configuration may also be implemented by (vice). The processing unit 202 is operated by the processor By interpreting and executing instructions from various programs, various data processing and programs Control can be performed. The programs that can be executed by the processor are those that the processor possesses. It is stored in at least one of the memory area and the storage unit 203.

[0123] The processing unit 202 may have main memory. Main memory may be volatile memory such as RAM. It has at least one of memory and non-volatile memory such as ROM.

[0124] For example, RAM can be DRAM (Dynamic Random Access Memory). memory), SRAM (Static Random Access Memory) These are used, and a memory space is virtually allocated and used as the workspace for the processing unit 202. The operating system and application programs stored in memory unit 203 are... The program modules, program data, and lookup tables are used during execution. These data, programs, and are loaded into RAM. Each program module is directly accessed and operated by the processing unit 202.

[0125] The ROM contains BIOS (Basic Input / Output) which does not require rewriting. It can store the ut System and firmware, etc. As for ROM, Mask ROM, OTPROM (One Time Programmable Read) Only Memory), EPROM (Erasable Programmability) Examples include e Read Only Memory. EPROMs are also available in ultraviolet light. UV-EPROM (Ultra-Violet) allows for the erasure of stored data by irradiation. Erasable Programmable Read Only Memory) , EEPROM (Electrically Erasable Programmable) Examples include Read Only Memory (LEM), flash memory, and others.

[0126] [Storage section 203] The memory unit 203 has the function of storing the program to be executed by the processing unit 202. The storage unit 203 stores, for example, the calculation results generated by the processing unit 202 and the input to the input unit 201. It may also have a function to store the data. Specifically, the storage unit 203 is a processing unit. The graph structure generated in 202 (for example, the graph structure 21 shown in Embodiment 1) is calculated. It is preferable that the system has a function to store the results of the similarity assessment, etc.

[0127] The storage unit 203 includes at least one of volatile memory and non-volatile memory. The memory unit 203 may have, for example, volatile memory such as DRAM or SRAM. Memory section 203 is, for example, ReRAM (Resistive Random Access). Memory (also called resistive random-access memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectri c Random Access Memory), MRAM (Magnetoresi (Also known as magnetically resistive random access memory.) , or it may have non-volatile memory such as flash memory. Also, storage unit 20 3 refers to hard disk drives (HDDs) and solids Storage media such as Solid State Drives (SSDs) It may have a drive.

[0128] [Database 204] The document search system 200 may have a database 204. For example, data Base 204 consists of multiple documents and multiple graph structures for each of those documents. It has a function to store the creation. For example, the multiple documents stored in database 204 As a target, a method for searching documents according to one aspect of the present invention may be used. Alternatively, a database may be used. A conceptual dictionary may be stored in cell 204.

[0129] Note that the memory unit 203 and the database 204 do not necessarily have to be separated from each other. The document search system 200 has the functions of both a storage unit 203 and a database 204. It may have a memory unit.

[0130] Note that the memories of the processing unit 202, the storage unit 203, and the database 204 can each be regarded as an example of a non-temporary computer-readable storage medium.

[0131] [Display unit 205] The display unit 205 has the function of displaying the calculation results in the processing unit 202. Also, the display unit 205 has the function of displaying the compared documents and the similarity results. Further, the display unit 205 may have the function of displaying the documents specified for search.

[0132] Note that the document search system 200 may have an output unit. The output unit has the function of supplying data to the outside.

[0133] [Transmission path 206] The transmission path 206 has the function of transmitting various data. The transmission and reception of data between the input unit 201, the processing unit 202, the storage unit 203, the database 204, and the display unit 205 can be performed via the transmission path 206. For example, data such as the document specified by the user for search and the graph structure for the document to be compared are transmitted and received via the transmission path 206.

[0134] [Configuration example 2 of document search system] FIG. 8 shows a block diagram of a document search system 210. The document search system 210 includes a server 220 and a terminal 230 (such as a personal computer).

[0135] The server 220 includes a processing unit 202, a transmission path 212, a storage unit 213, and a communication unit 217a. Although not shown in FIG. 8, the server 220 may further include an input / output unit and the like.

[0136] ​​​​​​​​​​​Terminal 230 consists of an input unit 201, a storage unit 203, a display unit 205, a transmission line 216, and a communication unit 2 It has 17b and a processing unit 218. Although not shown in Figure 8, terminal 230 further has They may have a database or similar.

[0137] The user of the document search system 210 enters the document into the input unit 201 of the terminal 230. This document is a document designated by the user for search purposes and corresponds to document 20 shown in Embodiment 1. The document is sent from the communication unit 217b of terminal 230 to the communication unit 217a of server 220. It is believed.

[0138] The document received by the communication unit 217a is stored in the storage unit 213 via the transmission line 212. Alternatively, the above document may be supplied directly from the communication unit 217a to the processing unit 202. stomach.

[0139] The creation of graph structures and the calculation of similarity described in Embodiment 1 do not require high processing power. The processing unit 202 of server 220 is compared to the processing unit 218 of terminal 230. All processing capabilities are high. Therefore, the creation of the graph structure and the calculation of similarity are performed by processing unit 20. It is preferable to carry this out in step 2.

[0140] Then, the processing unit 202 calculates the similarity. The similarity is transmitted via the transmission line 212. It is stored in the memory unit 213. Alternatively, the similarity score is sent directly from the processing unit 202 to the communication unit 217. It may be supplied to a. The similarity is the communication from the communication unit 217a of server 220 to terminal 230. The data is transmitted to unit 217b. The similarity score is displayed on the display unit 205 of terminal 230.

[0141] [Transmission lines 212 and 216] Transmission paths 212 and 216 have the function of transmitting data. The data transmission and reception among the processing unit 202, the memory unit 213, and the communication unit 217a can be performed via the transmission path 212. The data transmission and reception among the input unit 201, the memory unit 203, the display unit 205, the communication unit 217b, and the processing unit 218 can be performed via the transmission path 216.

[0142] [Processing unit 202 and processing unit 218] The processing unit 202 has the function of performing calculations using the data supplied from the memory unit 213, the communication unit 217a, etc. The processing unit 218 has the function of performing calculations using the data supplied from the input unit 201, the memory unit 203, the display unit 20 5, and the communication unit 217b, etc. For the processing unit 202 and the processing unit 218, reference can be made to the description of the processing unit 202. The processing unit 202 preferably has a higher processing capacity than the processing unit 218.

[0143] [Memory unit 203] The memory unit 203 has the function of storing the program executed by the processing unit 218. Also, the memory unit 203 has the function of storing the calculation results generated by the processing unit 218, the data input to the communication unit 217b , and the data input to the input unit 201, etc.

[0144] [Memory unit 213] The memory unit 213 has the function of storing a plurality of documents, the graph structure for each of the plurality of documents, the calculation results generated by the processing unit 20 2, and the data input to the communication unit 217a, etc.

[0145] [Communication unit 217a and communication unit 217b] Using the communication unit 217a and the communication unit 217b, data is transmitted between the server 220 and the terminal 230. ​It can send and receive data. Communication units 217a and 217b include a hub, Routers, modems, etc. can be used. Data can be transmitted and received using either a wired or wireless connection. For example, radio waves, infrared rays, etc. may be used.

[0146] Note that communication between server 220 and terminal 230 is via the World Wide Web (WWW). The foundation of ) is the internet, intranet, extranet, PAN (Pers onal Area Network), LAN(Local Area Network) k), CAN (Campus Area Network), MAN (Metropol Itan Area Network), WAN (Wide Area Network) ), GAN (Global Area Network), and other computer networks This can also be done by connecting to [the appropriate network].

[0147] This embodiment can be combined with other embodiments as appropriate. [Explanation of symbols]

[0148] :10: Multiple documents, 10_1: Document, 10_i: Document, 10_n: Document, 11: Multiple Graph structure, 11_1: Graph structure, 11_i: Graph structure, 11_n: Graph structure, 2 0: Document, 21: Graph structure, 100: Document search system, 101: Input section, 102: G Rough structure creation unit, 103: Similarity calculation unit, 104: Output unit, 105: Storage unit, 112: Approximate Dictionary, 200: Document search system, 201: Input unit, 202: Processing unit, 203: Memory unit 204: Database, 205: Display unit, 206: Transmission line, 210: Document search system , 212: transmission line, 213: memory unit, 216: transmission line, 217a: communication unit, 217b: communication Information Unit, 218: Processing Unit, 220: Server, 230: Terminal

Claims

[Claim 1] It has an input section, a first processing section, a storage section, a second processing section, and an output section. The input unit has the function of inputting the first document to the first processing unit. The first processing unit has the function of creating a first graph structure from the first document. The first graph structure has a first node, a second node, and an edge. The edge has a first direction, The first direction is determined based on the semantic relationship between the first node and the second node, The storage unit has the function of storing the second graph structure, The second processing unit has a function to calculate the similarity between the first graph structure and the second graph structure. The output unit has the function of supplying information regarding the calculated similarity. Document search system.