Entailment search system and entailment search method

The implicature search system efficiently performs entailment analysis on large document sets by structuring documents into chapters and sections, searching for similar text fragments, and determining implications, thus enhancing document search alignment with user intent.

JP7867897B2Active Publication Date: 2026-06-01HITACHI LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2022-07-13
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

Conventional entailment analysis methods face computational challenges when applied to large volumes of documents, making it difficult to efficiently perform entailment analysis that aligns with user intent.

Method used

An implicature search system that includes a first input unit, a second input unit, a chapter and section structure analysis unit, a pre-search unit, and an implication determination unit to analyze document structure, search for similar text fragments, and determine implications based on user intent.

Benefits of technology

Enables efficient entailment analysis on large volumes of documents, aligning with user intentions and improving document search accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007867897000001
    Figure 0007867897000001
  • Figure 0007867897000002
    Figure 0007867897000002
  • Figure 0007867897000003
    Figure 0007867897000003
Patent Text Reader

Abstract

To provide an implication search system capable of achieving document research according to the intention of a user by efficiently performing implication analysis on large amounts of documents.SOLUTION: An implication search system 100 is configured including: a first input unit 101 that accepts documents; a second input unit 102 that accepts input intention; a chapter structure analysis unit 103 that analyzes the chapter structure of a document and generates text fragments; a preliminary search unit 110 that searches for text fragments that show similarity to the input intention; an implication determination unit 105 that performs implication determination based on the search results and the input intention; and an output part 106 that outputs the result of the implication determination.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an implicature retrieval system and an implicature retrieval method.

Background Art

[0002] For a certain statement, the fact that the meaning of that statement is included in another statement is called an implicature. For example, the statement "I play soccer" is included in the meaning of the statement (hypothesis) "I exercise", so it is said to implicate the hypothesis. The technique of analyzing the implicature relationship between such statements is called implicature analysis.

[0003] For example, Non-Patent Document 1 discloses a high-precision implicature analysis technique using a pre-trained language model based on deep learning and an example of item checking in a confidentiality agreement using the same. By using implicature analysis in this way, high-performance document investigation can be mechanically performed.

[0004] For example, in the document investigation means relying on search terms as shown in Patent Document 1, by using a "search term", which is one form representing the search intention of the user, to extract a sentence containing the search term, a document containing the user's intention can be investigated. On the other hand, when using implicature analysis, the user's intention can be given in the form of a sentence or the like, and it is not necessary to explicitly include a "search term" in the investigation target.

[0005] Therefore, using the above example of implicature, a document describing "I play soccer" can be extracted from the statement "I exercise" without including the word "soccer" in the search term.

[0006] Also, in a search engine using search terms, in order to express a complex argument relationship, it is necessary to use an advanced search formula, but in implicature analysis, such a complex argument relationship can also be handled in the form of a natural sentence.

[0007] Furthermore, depending on the type of implicature analysis, there are some that can distinguish "contradiction" and "non-reference" in addition to "implicature".

[0008] For example, the hypothesis "I don't play soccer" contradicts the statement "I play soccer," while the hypothesis "I play music" makes no mention of the statement "I play soccer."

[0009] In particular, sentences containing the word "contradiction" are highly likely to contain the search term, and excluding contradictory statements from the implication is important in a detailed examination of the document.

[0010] For example, such meticulous document analysis is necessary for investigations to determine whether legal documents violate laws or precedents, audits to check for fraudulent activities within an organization, and investigations to extract desired information from large volumes of investigation reports. In such cases, the number of documents to be investigated is often enormous, making machine-assisted investigation desirable. [Prior art documents] [Patent Documents]

[0011] [Patent Document 1] WO2020 / 079752 [Non-patent literature]

[0012] [Non-Patent Document 1] Yuta Koreeda and Christopher Manning, ContractNLI: A Dataset [Overview of the project] [Problems that the invention aims to solve]

[0013] Thus, performing entailment analysis as described in Non-Patent Document 1 is effective in document research. However, performing entailment analysis on a large number of documents, as is done with conventional search engines, has been difficult from a computational load standpoint.

[0014] This is because entailment analysis can only be performed when two descriptions of the subject to analysis are provided, and in document research, one of these descriptions is entered as the user's research intention. Therefore, entailment analysis cannot be performed until that input is made.

[0015] Furthermore, since entailment analysis using deep learning is computationally more demanding than conventional search methods, it is difficult to perform calculations for all documents in a large volume once the user's research intent has been input.

[0016] Therefore, the objective of the present invention is to provide a technology that enables efficient entailment analysis of a large volume of documents and realizes document research that aligns with user intent. [Means for solving the problem]

[0017] The implication search system of the present invention, which solves the above problems, is an implication search system that determines whether an input intent is implication in a document, and is characterized by including: a first input unit that receives the document; a second input unit that receives the input intent; a chapter and section structure analysis unit that analyzes the chapter and section structure of the document and generates text fragments; a pre-search unit that searches for text fragments that show similarity to the input intent; an implication determination unit that performs an implication determination based on the search results and the input intent; and an output unit that outputs the result of the implication determination.

[0018] Furthermore, the implication search method of the present invention is characterized in that, when an information processing system determines whether an input intent is imposed on a document, it performs the following steps: receiving the document; receiving the input intent; analyzing the chapter and section structure of the document to generate text fragments; searching for text fragments that show similarity to the input intent; performing an implication determination based on the search results and the input intent; and outputting the results of the implication determination. [Effects of the Invention]

[0019] According to the present invention, it is possible to efficiently perform implicature analysis on a large number of documents and realize document search that conforms to user intentions.

Brief Description of Drawings

[0020] [Figure 1] It is a block diagram showing a functional configuration example of the implicature search system in this embodiment. [Figure 2] It is a block diagram showing a hardware configuration example of a computer device that realizes the implicature search system of this embodiment. [Figure 3] It is a flowchart of the implicature search method in this embodiment. [Figure 4] It is a schematic diagram showing an example of text fragmentation processing in the implicature search method of this embodiment. [Figure 5] It is a flowchart of the implicature search method in this embodiment. [Figure 6] It is a schematic diagram showing an example of the user interface of the implicature search system of this embodiment.

Modes for Carrying Out the Invention

[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In the following, for the same or similar elements and processes, the same reference numerals are given and differences are described, and redundant descriptions are omitted. Also, for the embodiments described later, differences from the already described embodiments are described, and redundant descriptions are omitted. Further, each embodiment and its modifications can be combined in part or in whole within the scope consistent with the spirit of the present invention.

[0022] Figure 1 is a block diagram showing an example of the functional configuration of the implication search system 100 in this embodiment. The implication search system 100 delivered in this embodiment consists of an input unit 101 for registering documents to be searched in advance, a chapter and section structure analysis unit 103 for analyzing the chapter and section structure of the document and registering it in a pre-search unit 110 which is a database, an input unit 102 for receiving input of words, sentences, and phrases (hereinafter referred to as queries) given by the user as their search intent, a query processing unit 104 for converting the query into an appropriate format, an implication determination unit 105 for performing an implication determination on the search results in the pre-search unit 110, and an output unit 106 for displaying and presenting the results of the implication determination to the user.

[0023] First, the hardware configuration of the implication search system 100 will be described. The implication search system 100 can be implemented using a computer device.

[0024] Figure 2 is a block diagram showing an example of the hardware configuration of a computer device 10 that implements the entailment search system 100.

[0025] As shown in Figure 2, the computer device 10 is comprised of a processor 11, a storage device 12, an input device 13, an output device 14, and a communication interface 15, with each component connected to the others by a bus 16.

[0026] Of these, processor 11 can be assumed to be a CPU (Central Processing Unit) that has the function of controlling the computer device 10.

[0027] Furthermore, the memory device 12 is a memory device composed of a non-volatile memory device or a volatile memory device that stores programs and data for implementing the necessary functions, and serves as a work area for the processor 11.

[0028] The specific examples of the memory device 12 are not limited to those mentioned above; for example, they could be ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), or SSD. Flash memory such as (Solid State Drive) can be used.

[0029] Furthermore, the processor 11 and the memory device 12 are GPUs (Graphical Processing Units). A device using ) is also acceptable.

[0030] Specifically, for example, each processing unit of the implication search system 100 shown in Figure 1 (input units 101, 102, chapter / section structure analysis unit 103, query processing unit 104, implication determination unit 105, output unit 106, pre-search unit 110, embedded representation generation unit 111) is realized by the processor 11 executing a temporary or non-temporary program stored in the memory device 12.

[0031] Furthermore, the input data handled by the entailment search system 100 is stored, for example, in the storage device 12. In addition, various types of data held in the pre-search unit 110 shown in Figure 1, which will be described later, are also stored, for example, in the storage device 12.

[0032] The processor 11 consists of one or more processing units. 1 may include one or more arithmetic units and multiple processing cores. Processor 11 may be implemented as one or more central processing units, microprocessors, digital signal processors, microcontrollers, microcomputers, state machines, logic circuits, graphics processing units, chip-on systems, or any device that performs signal manipulation by control instructions, etc.

[0033] In a computer device 10 that implements the entailment search system 100, the program executed by the processor 11 may include an OS (Operating System).

[0034] Furthermore, the program executed by the processor 11 may include various programs for realizing the functions of each processing unit of the implication search system 100 (for example, input programs for input units 101 and 102, chapter and section structure analysis programs for chapter and section structure analysis 103, query processing programs for query processing unit 104, pre-search processing programs for pre-search unit 110, implication determination programs for implication determination unit 105, and output programs for output unit 106).

[0035] By executing and operating the programs described above, the processor 11 can function as an input unit 101, 102, a chapter / section structure analysis unit 103, a query processing unit 104, and a pre-search unit 110.

[0036] In the computer device 10 shown in Figure 2, software elements such as the OS and various programs are stored in one of the storage areas of the storage device 12. The OS and various programs may be pre-recorded on a portable recording medium, in which case the programs are read from the portable recording medium by a media reader and stored in the storage device 12. Alternatively, the OS and various programs may be acquired via a communication medium.

[0037] The input device 13 is a device that allows the user to execute commands and data input to the entailment search system 100, and is specifically implemented as, for example, a mouse, keyboard, touch panel, microphone, or scanner.

[0038] The output device 14 is a device that performs data output from the entailment search system 100, and is specifically implemented as, for example, a display, printer, or speaker.

[0039] The communication interface 15 is a device that connects to the external network of the computer device 10 and transmits and receives various types of data handled by the entailment search system 100, and is specifically implemented by, for example, a NIC (Network Interface Card).

[0040] When the implication search system 100 is implemented on a computer device 10 equipped with a communication interface 15, the implication search system 100 can be configured to send and receive data from another terminal via an external network.

[0041] Furthermore, the entailment search system 100 is not limited to a configuration implemented on a single computer (computer device) as shown in Figure 2 (computer device 10), but may also be implemented on a computer system consisting of multiple computers (computer devices).

[0042] In that case, the computers can communicate with each other via a network, and for example, multiple functions of a language model processing unit may be implemented on multiple computers.

[0043] This concludes the explanation of the hardware configuration of the entailment search system 100. Now, we will return to the explanation of the functional configuration of the entailment search analysis system 100 shown in Figure 1.

[0044] The implication search system 100 has two basic processing functions: a process for receiving and registering documents to be investigated, and a search process for investigating portions of the registered documents that are implication of the user's intent.

[0045] Figure 3 shows the registration process flow in the entailment search system 100. The configuration of the registration function of the entailment search system 100 will be explained according to each processing step in the flow shown in Figure 3.

[0046] First, the input unit 101, which is involved in the information extraction process (301) from the input document, will be described. The document input by the user is assumed to be in text format. Here, the text format may be PlainText consisting only of so-called words, or it may be a format such as HTML, XML or JSON that includes some kind of structural information.

[0047] Furthermore, document images can also be used as input, but in that case, the document image must be processed using OCR or similar methods. Input should be in text format. Images, audio, and other information embedded in the text may also be included, but information that represents the content of the document must be in text format.

[0048] The input unit 101 extracts information about the document in a specific format according to the data format and input information of the document entered by the user. This information extraction refers to, for example, storing the information in a specific format on memory or storage devices within the hardware configuration.

[0049] For example, if the document is PlainText, line breaks and general sentences within the text contained in the document will be lost. By applying segmentation techniques and other methods, information is extracted as a collection of sentences or phrases (a collection of text fragments).

[0050] Furthermore, if the document is XML containing layout information, the document is usually of a certain nature. It stores text boxes, which consist of coherent sentences, and their location information on the paper or display device. In such cases, the string information within each text box is treated as text fragments, and the information is extracted in a format that associates their location information with each fragment.

[0051] Furthermore, if the document contains semantic structure information such as HTML, then, similar to XML which includes layout information, each string of information is extracted as a text fragment, and the associated structure information is also extracted. Information is extracted in a format that associates the relevant elements.

[0052] In this way, the input unit 101 extracts information in a specific format and transmits the extracted information, i.e., document information, to the chapter and section structure analysis unit 103. The chapter and section structure analysis unit 103 performs a process (302) to analyze the chapter and section structure within the document.

[0053] The chapter / section structure analysis unit 103 has the function of reconstructing the document information extracted by the input unit 101 into a meaning-based structure, thereby providing a semantic tree structure to the target document.

[0054] For example, if the input document is a patent document, then a single patent document (the root of the structural tree) has four sub-documents: specification, claims, abstract, and drawings. Of these, the specification, for instance, has further sub-structures such as technical field and background art. The background art then has prior art documents, and the prior art documents have an information structure of patent and non-patent documents as its children.

[0055] Chapter and section structure analysis is based on this tree structure and involves analyzing the text information that corresponds to each section. This involves estimating the relationship between elements. In this invention, it is used to clarify the semantic relationships between the aforementioned text fragments.

[0056] The chapter / section structure analysis unit 103 may use the structure extracted by the input unit 101 as is if the document input by the input unit 101 holds a semantic structure in a tree structure, such as in HTML format.

[0057] Also, for example, if the entered document contains data such as PlainText or layout information, Based on layout position information, font information, and features extracted from the text contained in the fragments, semantic relationships can be estimated using rules or machine learning methods.

[0058] When implemented according to the above rules, for example, in the case of a horizontally written Japanese document, if layout information is used, text fragments located in the upper left corner or represented by a larger font than the surrounding text are determined to be higher-level elements representing chapter / section names, and the text fragments directly below those chapter / section text fragments are determined to be the content of the chapter / section.

[0059] Furthermore, when using text information, text fragments that begin with numbers or those containing headings such as chapters and sections relatively close to the beginning can be treated as text fragments representing chapter and section names, while others can be processed in the same way as when using layout information. Of course, it is also possible to use both layout information and text information for the determination.

[0060] Furthermore, when using machine learning methods, text fragments such as chapter and section names and text fragments representing their associated content can be distinguished in advance, and machine learning can be performed using this distinguished data. In this case, the features of the machine learning model may be the same text information as when using layout information or rules, or, for example, feature extraction from text using a pre-trained deep learning model.

[0061] The machine learning classification model may simply use a tri-classifiable classifier that is independent of the hierarchy of text fragments, or it may apply graph-based methods or transition-based methods used in parsers to the text fragments.

[0062] In this way, the chapter-section structure analysis unit 103 can clarify the relationships between text fragments. Therefore, based on these relationships, the chapter-section structure analysis unit 103 generates text fragments such that each text fragment includes the higher-level text fragment. (303) Figure 4 shows an overview of the text fragment generation process, including the higher-level structure. For example, if the target document is a patent document as shown in Figure 4, the text fragment will include "background art" if the item is, for example, "patent document".

[0063] By structuring text fragments in this way, entailment analysis can be performed efficiently during the search process described later.

[0064] In this manner, the chapter / section structure analysis unit 103 generates one or more text fragments for a single document and transmits the generated text fragments to the pre-search unit 110.

[0065] Meanwhile, the pre-search unit 110 performs processes (304) and (305) to complete the registration process.

[0066] The pre-search unit 110 comprises an embedded representation generation unit 111. An embedded representation is a set of features that have vector or tensor form and characterize text fragments. However, assuming that an embedded representation is mainly a set of features that have vector form, and insofar as various similar operations can be performed, The data format is not restricted.

[0067] An embedded representation is generated for each text fragment in process (304). The function of the embedded representation in this invention is to enable the estimation of similarity between text fragments through calculations between the embedded representations.

[0068] The means for calculating these embedded representations and the calculations used to indicate similarity will be described in detail later in the search process section.

[0069] The pre-search unit 110 stores the aforementioned text fragments and at least their corresponding embedded representations as a single record (search unit). Furthermore, a data storage mechanism similar to that of conventional search engines can be used to store the text fragments. In addition, other information provided during data input can also be stored in each record.

[0070] Next, let's discuss the search process. There are two main differences between entailment analysis and conventional text-similarity-based searching. First, the lexical overlap between the text expressing the user's research intent (query) and the entailing sentences in the actual document is not always large. Second, even documents similar to the query may, upon closer examination, contain content that expresses a "contradiction" rather than an entailment.

[0071] The second difference mentioned above can be rephrased as the presence of both implicational and contradictory texts within similar text sets. Therefore, by collecting texts similar to the query, it is possible to present documents that are either implicational, contradictory, or both, and that are suitable for the user's research intent.

[0072] For example, if you want to investigate whether a computer program code can be disclosed among a large number of confidentiality agreements, the user's investigation intent is "disclosure of computer program code," but the investigation can be useful for both implications (i.e., agreements that allow disclosure) and contradictions (i.e., agreements that do not allow disclosure).

[0073] On the other hand, if you only want to search for "contracts that allow the disclosure of computer programs," the query is similar, but the search will only include those that imply such contracts.

[0074] In either case, the goal is to collect contracts that mention the disclosure of computer programs. In this function, it is preferable to select documents from the collection based on their similarity to the query, and then further select them based on whether they have implications.

[0075] The search process in this embodiment broadly implements the process described above. Here, the search process will be explained using the search process flowcharts shown in Figures 1 and 5.

[0076] Here, the user's research intent (query) is taken as input. The input unit 102 can accept such research intent input. The query includes at least a string representing the research intent, and may also include conditions regarding implication relationships necessary for the research, such as "implication" or "contradiction." Furthermore, queries may generally include keywords or phrases that must be included, as is done in typical text searches, but for the sake of simplicity, here we will refer to the string representing the research intent and the information with implication relationship conditions specific to the implication search system 100 as a query. The string representing the research intent is a sentence, phrase, or word that expresses the research intent in natural language.

[0077] The query processing unit 104 processes information for such filtering, strings representing the research intent, It can extract and store implicational relationship conditions such as "implication" and "contradiction" from user input.

[0078] When a query is provided by the user via the query processing unit 104, the embedded representation generation unit 111 in the pre-search unit 110 first calculates an embedded representation for the string contained in the query in order to investigate similar text fragments (501).

[0079] Here, to account for cases where there is not necessarily lexical overlap between the query and the document, it is desirable that the method for calculating the embedded representation vectors be based on fixed-dimension vectors, where synonyms and similar words are represented by similar vectors, rather than vectors where each word is assigned a dimension of the vector, as in the so-called Bag-of-Words model.

[0080] For example, as a word embedding representation, there are known methods such as word2vec, which represent words as points (vectors) of a fixed dimension (e.g., 300 dimensions).

[0081] The vectors provided by this word2vec are trained to maximize the dot product between the vector corresponding to the word of interest and the vectors corresponding to words that appeared within the same context (a certain range of words within the word of interest) during prior training. Due to this characteristic, words with similar contextual words will have similar vectors as synonyms.

[0082] By using word2vec in this way, synonyms that do not appear in the query can be handled efficiently. To calculate an embedded representation for any specific text (including both strings and text fragments contained in the query), first, morphological analysis is applied to the text, vectors corresponding to each morpheme are extracted from pre-calculated sources such as word2vec, and the sum of these extracted vectors is taken.

[0083] When calculating the sum as described above, words that are not related to meaning (so-called stop words) may be removed, or a weighted sum may be used instead of a sum. In this case, the weights may be determined by the importance of the words estimated from their frequency of appearance, or by so-called machine learning methods.

[0084] Furthermore, not limited to word2vec, any embedding representation can be used whose objective function is to express the similarity between a word and a context word using methods such as the dot product.

[0085] Furthermore, embedded representations using deep learning models (pre-trained language models) such as recurrent neural networks and Transformers can be used. These pre-trained language models, like word2vec, are trained using a large amount of text data beforehand to predict words and phrases from a given context (sentences, missing words, phrases, words), and, like word2vec, produce learning results that represent the similarity of words and phrases as the closeness of vectors in a fixed dimension based on contextual information. Here, closeness refers to measures assumed during the pre-training, such as dot product or distance.

[0086] Generally, such pre-trained language models can generate vectors that represent sentences and phrases, and these can be used as embedding representations for any text.

[0087] Furthermore, as in the example using word2vec mentioned above, the sum of vectors for each word can be used in a similar manner.

[0088] Furthermore, if the entailment analysis described later is implemented using a similar deep learning model, the vector calculated by the entailment determination unit 105 can also be used as the embedding representation.

[0089] When the embedded expression generation unit 111 provides an embedded expression to the query, the pre-search unit 110 performs a similar text search (ranking) using the embedded expression (502).

[0090] Each record stored in the pre-search unit 110 contains string information representing a text fragment, supplementary information given during the registration process, and an embedded representation. Records are ranked based on a similarity score (real value) with at least the embedded representation generated for the query.

[0091] In this case, suitable similarity scores include the inner product or cosine similarity between the embedded representation for the query and the embedded representation associated with the record. It is natural to use the same similarity measure used for pre-training the embedded representations.

[0092] However, when using embedded representations for words, for example, it is preferable to use the sum of the embedded representations of the words contained in a text fragment for a text fragment. In this case, the length of the embedded representation vector for the text fragment (usually the square root of its own vector dot product) is extended by the amount of the word. Taking this extension into account, it is also preferable to use, for example, cosine similarity (the cosine of the angle between two vectors).

[0093] Furthermore, if there is information appended to the record or filtering information attached to the query, the presence or absence of such information can be replaced with 0, 1, or an appropriate real number, and an overall similarity score can be calculated by summing or multiplying the similarity score between that number and the embedded representation.

[0094] For example, if the accompanying information is an AND condition that it must contain a certain word, it is preferable to use the product of the similarity score between the embedded representations, assigning 1 if it contains the word and 0 if it does not.

[0095] Such calculations can be performed in the same way as with existing text search engines, and can also be provided by the user as information accompanying the query.

[0096] For example, when using the dot product or cosine similarity as the similarity score, a higher score indicates greater similarity. The polarity of this score can be freely reversed by determining the sign of the score or by defining the quotient relative to 1 as a new score. Therefore, here we will explain it as a case where a higher score indicates greater similarity, without losing generality.

[0097] Now, by calculating an overall similarity score, the records stored in the pre-search unit 110 can be ranked in descending order of score. The pre-search unit 110 extracts text fragments from the top-ranked records and sends them to the entailment determination unit 105 along with the query and the information used to calculate the similarity score.

[0098] The implication determination unit 105 performs implication analysis on the text fragments transmitted from the pre-search unit 110 (503).

[0099] Although various methods exist for this implication analysis, the method using Transformers is preferable for performing implication analysis with high accuracy.

[0100] When using a pre-trained language model with Transformer for implication analysis, the string representing the user's research intent included in the query is considered a "hypothesis" in the implication, and each text fragment is used as a description to determine whether or not it implies something.

[0101] The input is a string formed by combining the above hypothesis and description with a special token. For example, if the hypothesis is "The technical field of the invention is a search using implication", the special token is "[SEP]", and the description is a text fragment containing the technical field generated from this specification ("Specification", "Technical Field", "The present invention is..."), then the input text would be "The technical field of the invention is a search using implication [SEP] "Specification", "Technical Field", "The present invention is...".

[0102] Incidentally, when performing the implication determination described above, it is not always obvious that the description of the technical field is a reference to the "technical field." However, since the higher-level chapter-section structure explicitly states "technical field," using this as input for implication determination demonstrates that the chapter-section structure analysis effectively manages information related to the user's research intent.

[0103] Furthermore, when obtaining an embedded representation of the entire input, additional special tokens are added to the beginning or end.

[0104] This is then divided into morphemes using a suitable morphological analyzer (usually the morphological analyzer included with the pre-trained language model being used). Here, morphemes are not linguistic morphemes, but rather tokens pre-registered in the pre-trained language model. In this case, it is preferable to use the special tokens registered in the pre-trained language model.

[0105] A pre-trained language model can assign an embedding representation to each morpheme. For the embedding representation of the entire input, operations such as summation, attention mechanisms, average values, and maximum values ​​of each dimension of a vector may be calculated for each morpheme's embedding representation, or the embedding representation assigned to the special token representing the entire input may be used.

[0106] The embedded representation of the obtained input can be used to identify implication labels using any classifier (generally a single perceptron is commonly used). A score can then be generated for each implication label as a result of this identification.

[0107] In training this classifier, a pre-trained language model may be retrained using pre-created entailment data.

[0108] Furthermore, while implication labels only need to be able to distinguish between "implication" and other types of statements, it is preferable for them to have at least three labels: "implication," "contradiction," and "non-reference."

[0109] When using an implication analysis model as the embedded expression generation unit 111, it is not possible to obtain a "hypothesis" because the user's research intent is not provided during document registration. For this reason, the "hypothesis" field can be left blank, or the text fragment itself can be provided as the "hypothesis." It is preferable to use the embedded expression of the entire input as described above for the generated embedded expression.

[0110] The implication determination unit 105 calculates an overall score by replacing the score for the implication relationship label with the similarity score between the embedded expressions, ranks the text fragments in descending order of score, and sends them to the output unit 106.

[0111] In this case, if the user's research intent is solely "implication," the "implication" score can be used. If the user's research intent is either "implication" or "contradiction," and the implication relationship label includes at least the three values ​​mentioned above, it is preferable to use a score that is not "unmentioned" (the score with the polarity of the unmentioned score reversed).

[0112] The output unit 106 outputs the survey results to the user. Figure 6 shows the entailment search system. This shows an example of a user interface for input and output in M100.

[0113] The user interface exemplified here preferably includes, as input, a means 601 for inputting the investigation intent, and means (602 and 603) for indicating that either implication, contradiction, or both are the subject of the investigation.

[0114] When the user enters the research intent and clicks the "Research" button 605 to initiate the research, the entailment search system 100 presents the research results in the research results table 604.

[0115] The results of the investigation are presented in the order ranked by the entailment search system 100, showing text fragments and their associated documents or document IDs.

[0116] Such a user interface can present the user with an overall score 606 derived from implication analysis, as well as the results 607 of implication relationship labels. The implication relationship labels, in particular, are meaningful in interpreting the survey results, and their presentation is preferable.

[0117] In this way, it becomes possible to efficiently calculate and present entailment analysis results to users from a large volume of documents, according to the user's research intent.

[0118] The embodiments described above are detailed for the purpose of clearly illustrating the present invention and are not necessarily limited to those comprising all the elements and configurations described. Therefore, the present invention is not limited to the embodiments described above and may include various modifications within a reasonable scope. For example, as long as it does not contradict the present invention, some elements or configurations of one embodiment may be replaced with configurations of other embodiments, and elements or configurations of other embodiments may be added to the elements or configurations of one embodiment. Furthermore, some elements or configurations of each embodiment may be added, deleted, replaced, integrated, or distributed. In addition, the elements, configurations, and processes shown in the embodiments may be distributed, integrated, or rearranged as appropriate based on processing efficiency or implementation efficiency.

[0119] Furthermore, each of the above configurations, functions, processing units, processing means, etc., may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the above configurations, functions, etc., may be implemented in software by a processor interpreting and executing programs that implement each function. Information such as programs, tables, and files that implement each function is stored in memory, a hard disk, an SSD (Solid State Drive), or other storage devices. It can be stored on recording media such as IC cards, SD cards, and DVDs.

[0120] Furthermore, the control lines and information lines shown in the drawings are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it can be assumed that almost all components are interconnected.

[0121] According to this embodiment, it becomes possible to efficiently perform entailment analysis on a large volume of documents and realize document research that aligns with the user's intent.

[0122] The following will be made clear from the description herein: that, in the implication search system of this embodiment, the pre-search unit comprises an embedded expression generation unit that generates embedded expressions with respect to the text fragment and the input intent for the purpose of calculating the similarity, and performs the similarity search based on the embedded expressions. This allows for more accurate and efficient calculation of similarity scores and other processes, resulting in generally better document research efficiency. Consequently, it becomes possible to perform entailment analysis more efficiently on large volumes of documents, enabling document research that aligns with user intent.

[0123] Furthermore, in the implication search system of this embodiment, the embedded expression generation unit may use the embedded expression used in the implication determination unit.

[0124] This will streamline the calculation of similarity scores in implication determination, and consequently enable more efficient implication analysis of large volumes of documents, thereby realizing document research that aligns with user intent.

[0125] Furthermore, in the implication search system of this embodiment, the chapter-section structure analysis unit may generate text fragments with higher-level semantic structures added.

[0126] This allows for efficient entailment analysis during search processing. Consequently, it enables more efficient entailment analysis of large volumes of documents, resulting in document research that aligns with user intent.

[0127] Furthermore, in the implication search method of this embodiment, the information processing system may further perform a process to generate embedded expressions with respect to the text fragment and the input intent in order to calculate the similarity during the search, and perform the similarity search based on the embedded expressions.

[0128] Furthermore, in the implication search method of this embodiment, the information processing system may use the embedded expression used for implication determination when generating the embedded expression.

[0129] Furthermore, in the implication search method of this embodiment, the information processing system may generate text fragments in which a higher-level semantic structure is added when analyzing the chapter and section structure. [Explanation of Symbols]

[0130] 10 Computer equipment 11 processors 12. Storage Devices 13 Input Devices 14 Output Devices 15 Communication Interface 16 bus 100 Entailment Search System 101 Input Section (Document Input) 102 Input section (query input) 103 Chapter and section structure analysis part 104 Query Processing Unit 105 Implication judgment part 106 Output section 110 Pre-search section 111 Embedded Expression Generation Unit 601 Survey Intent Input Section 602 Implication / Intent Input Section 603 Implication / Relationship Intent Input Section 604 Survey Results Output Unit

Claims

1. An implication search system for determining whether input intent is implied or contradictory for documents with the same or different data formats, A first input unit that receives the aforementioned document, A second input unit that receives the aforementioned input intent, A chapter-section structure analysis unit analyzes the chapter-section structure of the aforementioned document and generates multiple text fragments whose semantic relationships with each other are clear, A pre-search unit that searches for text fragments that show similarity to the input intent, An implication determination unit that performs implication determination based on the results of the search and the input intent, An output unit that outputs the result of the implication determination, An entailment search system characterized by including [this].

2. The pre-search unit comprises an embedded expression generation unit that generates embedded expressions with respect to the text fragment and the input intent in order to calculate the similarity, and performs the similarity search based on the embedded expressions. The implication search system according to feature 1.

3. The embedded expression generation unit uses the embedded expression used in the implication determination unit. The implication search system according to feature 2.

4. The aforementioned chapter and section structure analysis unit generates text fragments with the higher-level semantic structure added. 、 The implication search system according to feature 1.

5. Information processing system, When determining whether input intent is implied or contradictory for documents with the same or different data formats, The process of receiving the aforementioned document, The process of receiving the aforementioned input intent, A process to analyze the chapter and section structure of the aforementioned document and generate multiple text fragments whose semantic relationships with each other are clear, A process for searching for text fragments that show similarity to the input intent. A process for performing implication determination based on the results of the search and the input intent, The process of outputting the result of the implication determination, A method for entailment search characterized by performing the following.

6. The aforementioned information processing system In the aforementioned search, a process is further performed to generate embedded expressions with respect to the text fragment and the input intent in order to calculate the similarity, and the search is performed based on the similarity using the embedded expressions. The implication search method according to feature 5.

7. The aforementioned information processing system In generating the aforementioned embedded expression, the embedded expression used for the implication determination is used, The implication search method according to feature 6.

8. The aforementioned information processing system In analyzing the aforementioned chapter-section structure, text fragments are generated with the higher-level semantic structure added. The implication search method according to feature 5.