Natural language processing device

The natural language processing device addresses the challenge of considering context in document retrieval by using a teacher data generation unit and inference unit to accurately match requirements in higher-level documents with relevant lower-level documents.

WO2025134331A1PCT designated stage expired Publication Date: 2025-06-26NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/046002
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing natural language processing methods struggle to consider the context of chapters, sections, and items when searching for lower-level documents that satisfy requirements described in higher-level documents.

Method used

A natural language processing device is developed, comprising a teacher data generation unit that creates labels indicating the correspondence between chapters, sections, and items in higher-level documents and lower-level documents, and an inference unit that uses a machine learning model to search for lower-level documents in the correct order and context.

Benefits of technology

This solution enables effective searching for lower-level documents that satisfy requirements while considering the context of chapters, sections, and items, thereby improving the accuracy and relevance of document retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023046002_26062025_PF_FP_ABST
    Figure JP2023046002_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This natural language processing device comprises a teaching data generation unit and an inference unit. The teaching data generation unit generates teaching data expressing correspondence relationships between lower-level documents and each of chapters, subchapters, and sections included in upper-level documents. The inference unit acquires a chapter and a subchapter of a first upper-level document, said chapter and subchapter including a section that serves as a search condition, and on the basis of a machine learning model that was trained using the teaching data, performs a search of the lower-level documents, in accordance with reference relationships between the lower-level documents and in the order of said chapter, said subchapter, and said section of the first upper-level document, and outputs a lower-level document corresponding to said section of the first upper-level document.
Need to check novelty before this filing date? Find Prior Art

Description

Natural Language Processing Unit

[0001] The present invention relates to a natural language processing device.

[0002] When there is a higher-level document that describes multiple requirements and a lower-level document that describes the means to satisfy each of the requirements described in the higher-level document, the correspondence between the requirements described in the higher-level document and the locations in the lower-level document that describe the means to satisfy those requirements must be confirmed.

[0003] Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova: "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding", Proceedings of NAACL HLT 2019, pages 4171-4186. Pretrained Japanese BERT models released.<URL:https: / / www.nlp.ecei.tohoku.ac.jp / newsrelease / 3284 / >

[0004] When searching for a part of a lower-level document that corresponds to a requirement written in a higher-level document, a method that uses machine learning to associate sentences in a document (see, for example, Non-Patent Documents 1 and 2) can be considered.

[0005] However, when using machine learning, when creating training data, the implication relationships between higher-level documents and lower-level documents are labeled by item (or clause), which poses a problem in that it is not possible to take into account the context of chapters and sections.

[0006] This invention has been made in light of the above circumstances, and its purpose is to provide a natural language processing device that can search for lower-level documents that meet requirements taking into account the context of the chapters, sections, and paragraphs of the higher-level document in which the requirements are stated.

[0007] A natural language processing apparatus according to one aspect of the present invention includes a training data generation unit and an inference unit. The training data generation unit generates training data representing the correspondence between each of the chapters, sections, and terms included in a higher-level document and a lower-level document. The inference unit acquires the chapters and sections of a first higher-level document that includes a term as a search condition, and, based on a machine learning model generated using the training data, searches the lower-level documents in the order of the chapter, section, and term of the first higher-level document according to the reference relationships between the lower-level documents, and outputs the lower-level document corresponding to the term of the first higher-level document.

[0008] According to the present invention, it is possible to provide a natural language processing device capable of searching for lower-level documents that satisfy requirements taking into consideration the context of the chapters, sections, and paragraphs of the higher-level document in which the requirements are written.

[0009] FIG. 1 is a block diagram showing the functional configuration of a natural language processing apparatus according to an embodiment. FIG. 2 is a diagram showing an example of chapters, sections, and terms in a higher-level document or a lower-level document according to an embodiment. FIG. 3 is a flowchart showing natural language processing in the natural language processing apparatus according to an embodiment. FIG. 4 is a flowchart showing a process of generating teacher data in a teacher data generation unit according to an embodiment. FIG. 5 is a diagram showing the process of generating teacher data according to an embodiment. FIG. 6 is a diagram showing a matrix representing the correspondence between chapters of a higher-level document and lower-level documents according to an embodiment. FIG. 7 is a diagram showing a matrix representing the correspondence between sections of a higher-level document and lower-level documents according to an embodiment. FIG. 8 is a diagram showing the correspondence between terms of a higher-level document and lower-level documents according to an embodiment. FIG. 9 is a flowchart showing the process of generating inference result data in an inference unit according to an embodiment. FIG. 10 is a diagram showing the process of generating inference result data according to an embodiment. FIG. 11 is a diagram showing the process of generating inference result data according to an embodiment from another perspective. FIG. 12 is a diagram showing an example of the hardware configuration of a natural language processing apparatus according to an embodiment.

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS In the following description, components having the same functions and configurations are denoted by the same reference numerals.

[0011] 1. Functional Configuration of the Embodiment The functional configuration of a natural language processing apparatus 10 of the embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the functional configuration of the natural language processing apparatus 10.

[0012] The natural language processing device 10 of the embodiment is a device that, when searching for a lower-level document that corresponds to (or is compatible with) a requirement item of a higher-level document, searches for and extracts the lower-level document while taking into account the context of the chapter, section, and paragraph of the higher-level document in which the requirement item is written.

[0013] Here, we will explain the terms "high-level document," "low-level document," "chapter," "section," and "item." A high-level document is, for example, a document that describes multiple requirements. A low-level document is, for example, a document that describes the means to satisfy each requirement described in the high-level document.

[0014] A higher-level document or a lower-level document includes at least one of a chapter, a section, and a paragraph. A chapter includes one or more sections. A section includes one or more paragraphs. For example, a chapter represents a large division into which the entire text written in a higher-level document or a lower-level document is divided. A section represents a division into which a chapter is divided. Furthermore, a paragraph represents a smaller division into which a section is divided. In this specification, a "chapter" represents a sentence or meaning contained in a chapter, a "section" represents a sentence or meaning contained in a paragraph, and a "paragraph" represents a sentence or meaning contained in a paragraph.

[0015] An example of the structure of chapters, sections, and paragraphs in a higher-level document or a lower-level document will be described with reference to Figure 2. Figure 2 is a diagram showing an example of chapters, sections, and paragraphs in a higher-level document or a lower-level document. The higher-level document 20 or the lower-level document 21 includes chapter 1 and chapter 2. Chapter 1 includes sections 1, 2, and 3. Section 1 includes sections 1, 2, and 3. Section 2 includes sections 1 and 2. Section 3 includes section 1. Chapter 2 includes sections 1 and 2. Section 1 includes sections 1 and 2. Section 2 includes section 1. Note that the structure and number of chapters, sections, and paragraphs in the higher-level document and the lower-level document are not limited to those described above.

[0016] 1, the natural language processing apparatus 10 includes a teacher data generation unit 11, a pre-learning generation unit 12, and an inference unit 13. The natural language processing apparatus 10 further includes teacher data 14 as data, and an entailment learning model 15 as a machine learning model (or a deep learning model).

[0017] The teacher data generation unit 11 receives a plurality of higher-level documents 20 and a plurality of lower-level documents 21 as input data. The teacher data generation unit 11 generates teacher data 14 using the plurality of higher-level documents 20 and the plurality of lower-level documents 21. The teacher data generation unit 11 determines a correspondence between each of the chapters, sections, and paragraphs included in the higher-level document 20 and the lower-level documents 21, and generates a label indicating whether or not a correspondence exists. That is, the teacher data generation unit 11 determines an implication relationship between each of the chapters, sections, and paragraphs of the higher-level document 20 and the lower-level documents 21, and generates a label indicating whether or not an implication relationship exists. Specifically, the teacher data generation unit 11 generates a label indicating whether or not an implication relationship exists between each of the sentences (or sentence meanings) included in the chapters, sections, and paragraphs included in the higher-level document 20 and the sentences (or sentence meanings) included in the lower-level document 21. The implication relationship is determined to hold when, for example, if the chapter of the higher-level document 20 is true, then the lower-level document 21 is also true.

[0018] The pre-learning generation unit 12 uses a machine learning algorithm that inputs the training data 14 generated by the training data generation unit 11 to learn the entailment relationships between each of the chapters, sections, and terms of the higher-level document 20 and the lower-level document 21, and generates an entailment relationship learning model 15 for determining the entailment relationships. The entailment relationship learning model 15 is a model that infers whether each of the chapters, sections, and terms of the higher-level document entails the lower-level document, and learns to make the inference.

[0019] The inference unit 13 acquires information about the chapters and sections of a higher-level document that includes a requirement item (or a clause), and extracts lower-level documents 21 that are implicated by the chapters, sections, and requirement items, in the order of chapter, section, and requirement item, based on the entailment learning model 15. The inference unit 13 also generates reference relationships between the multiple lower-level documents 21, in other words, establishes links between the multiple lower-level documents 21, in accordance with the entailment learning model 15 or another natural language processing model. The inference unit 13 searches for the lower-level documents 21 that include the requirement item as a search condition (or search key), in the order of chapter, section, and requirement item, and narrows down the lower-level documents 21 according to the reference relationships between the lower-level documents 21, thereby extracting lower-level documents that correspond to (or match) the requirement item. The inference unit 13 then outputs inference result data 16 that includes the extracted lower-level documents.

[0020] The machine learning algorithm used by the pre-training generation unit 12 and the natural language processing model used by the inference unit 13 can be, for example, BERT (Bidirectional Encoder Representations from Transformers) described in Non-Patent Document 1 and Non-Patent Document 2.

[0021] 2. Operation of the embodiment The operation of the natural language processing apparatus 10 of the embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the natural language processing executed by the natural language processing apparatus 10.

[0022] First, the training data generating unit 11 receives a plurality of higher-level documents 20 and a plurality of lower-level documents 21 (S1).

[0023] Next, the teacher data generating unit 11 generates teacher data 14 using the received upper-level document 20 and lower-level document 21 (S2). The process of generating the teacher data 14 will be described in detail later.

[0024] Next, the pre-learning generation unit 12 generates an entailment learning model 15 using the teacher data 14 generated by the teacher data generation unit 11 (S3).

[0025] Next, the inference unit 13 searches for lower-level documents corresponding to the requirement items based on the chapters, sections, and requirement items of the higher-level document including the requirement items as the search conditions, the implication learning model 15, and the reference relationships between the lower-level documents, and outputs inference result data 16 including the lower-level documents extracted by the search (S4). The process of generating the inference result data 16 will be described in detail later.

[0026] 2.1 Training Data Generation Process Next, the training data 14 generation process executed by the training data generation unit 11 will be described with reference to FIGS. 4 to 8. FIG. 4 is a flowchart showing the training data 14 generation process in the training data generation unit 11. FIG. 5 is a diagram showing the training data 14 generation process. FIGS. 6, 7, and 8 show training data generated by the training data generation unit 11. More specifically, FIG. 6 shows a matrix (or table) representing the correspondence between chapters of a higher-level document and lower-level documents. FIG. 7 shows a matrix (or table) representing the correspondence between sections of a higher-level document and lower-level documents. Furthermore, FIG. 8 shows a matrix (or table) representing the correspondence between sections of a higher-level document and lower-level documents. Here, it is assumed that multiple higher-level documents 20_1, 20_2, 20_3, ..., 20_m and multiple lower-level documents 21_1, 21_2, 21_3, ..., 21_n are prepared. m and n are integers equal to or greater than 2.

[0027] As described above, the teacher data generation unit 11 associates the lower-level documents 21_1 to 21_n with each of the chapters, sections, and paragraphs of the higher-level documents 20_1 to 20_m. That is, the teacher data generation unit 11 generates labels indicating the correspondence between each of the chapters, sections, and paragraphs of the higher-level documents 20_1 to 20_m and the lower-level documents 21_1 to 21_n. In other words, the teacher data generation unit 11 generates labels indicating whether or not an implication relationship exists between each of the chapters, sections, and paragraphs of the higher-level documents 20_1 to 20_m and the lower-level documents 21_1 to 21_n.

[0028] First, as shown in Figures 4 and 5, the training data generation unit 11 associates the chapters of the higher-level documents 20_1 to 20_m with the lower-level documents 21_1 to 21_n (S11). Depending on whether the chapter of the higher-level document implies the lower-level document, a label indicating whether the chapter of the higher-level document and the lower-level document correspond to each other is generated. That is, if a sentence (or meaning of the sentence) in the chapter of the higher-level document implies a sentence (or meaning of the sentence) in the lower-level document, a label indicating that an implication relationship exists between the chapter of the higher-level document and the lower-level document is generated. Conversely, if a sentence in the chapter of the higher-level document does not imply a sentence in the lower-level document, a label indicating that an implication relationship does not exist between the chapter of the higher-level document and the lower-level document is generated.

[0029] Specifically, the teacher data generation unit 11 determines whether or not chapter 1 of the higher-level document 20_1 and the lower-level document 21_1 correspond to each other. That is, the teacher data generation unit 11 determines whether or not chapter 1 of the higher-level document 20_1 implies the lower-level document 21_1. As shown in Fig. 6, if chapter 1 of the higher-level document 20_1 implies the lower-level document 21_1, the teacher data generation unit 11 generates a label 1 indicating that chapter 1 of the higher-level document 20_1 and the lower-level document 21_1 correspond to each other. If chapter 1 of the higher-level document 20_1 does not imply the lower-level document 21_1, the teacher data generation unit 11 generates a label 0 indicating that chapter 1 of the higher-level document 20_1 and the lower-level document 21_1 do not correspond to each other.

[0030] Next, the training data generation unit 11 determines whether or not Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2 correspond to each other. That is, it determines whether or not Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_2. As shown in Fig. 6, if Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_2, the training data generation unit 11 generates a label 1 indicating that Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2 correspond to each other. If Chapter 1 of the higher-level document 20_1 does not imply the lower-level document 21_2, the training data generation unit 11 generates a label 0 indicating that Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2 do not correspond to each other.

[0031] Similarly, the teacher data generation unit 11 determines whether or not there is a correspondence between chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_3 to 21_n. That is, it determines whether or not chapter 1 of the higher-level document 20_1 implies each of the lower-level documents 21_3 to 21_n. As shown in FIG. 6, if chapter 1 of the higher-level document 20_1 implies each of the lower-level documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_3 to 21_n are in a correspondence relationship. If chapter 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that there is no correspondence between chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_3 to 21_n.

[0032] Similarly, the teacher data generation unit 11 determines whether or not chapter 2 of the higher-level document 20_1 corresponds to each of the lower-level documents 21_1 to 21_n. That is, it determines whether or not chapter 2 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n. As shown in FIG. 6, if chapter 2 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that chapter 2 of the higher-level document 20_1 corresponds to each of the lower-level documents 21_1 to 21_n. If chapter 2 of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that chapter 2 of the higher-level document 20_1 does not correspond to each of the lower-level documents 21_1 to 21_n.

[0033] The same applies to each of the chapters 3 and onward of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. The training data generation unit 11 determines whether each of the chapters 3 and onward of the higher-level document 20_1 corresponds to each of the lower-level documents 21_1 to 21_n. If each of the chapters 3 and onward of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 1. If each of the chapters 3 and onward of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 0.

[0034] The same applies to each of the chapters of the upper document 20_2 and each of the lower documents 21_1 to 21_n. The training data generation unit 11 determines whether each of the chapters of the upper document 20_2 corresponds to each of the lower documents 21_1 to 21_n. If each of the chapters of the upper document 20_2 implies each of the lower documents 21_1 to 21_n, the training data generation unit 11 generates a label 1. If each of the chapters of the upper document 20_2 does not imply each of the lower documents 21_1 to 21_n, the training data generation unit 11 generates a label 0.

[0035] The same applies to each of the chapters of the higher-level documents 20_3 to 20_m and each of the lower-level documents 21_1 to 21_n. The training data generation unit 11 determines whether each of the chapters of the higher-level documents 20_3 to 20_m corresponds to each of the lower-level documents 21_1 to 21_n, and generates a label 1 or a label 0, respectively.

[0036] Next, the training data generation unit 11 associates the sections of the higher-level documents 20_1 to 20_m with the lower-level documents 21_1 to 21_n (S12), as shown in Figures 4 and 5. Depending on whether the section of the higher-level document implies the lower-level document, a label is generated indicating whether the section of the higher-level document and the lower-level document correspond to each other. That is, if a sentence (or meaning of the sentence) written in the section of the higher-level document implies a sentence (or meaning of the sentence) written in the lower-level document, a label is generated indicating that an implication relationship exists between the section of the higher-level document and the lower-level document. Conversely, if a sentence written in the section of the higher-level document does not imply a sentence written in the lower-level document, a label is generated indicating that an implication relationship does not exist between the section of the higher-level document and the lower-level document.

[0037] Specifically, the training data generation unit 11 determines whether or not there is a correspondence between Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_1. That is, it determines whether or not Section 1 of Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_1. As shown in FIG. 7 , if Section 1 of Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_1, the training data generation unit 11 generates a label 1 indicating that Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_1 are in a correspondence relationship. If Section 1 of Chapter 1 of the higher-level document 20_1 does not imply the lower-level document 21_1, the training data generation unit 11 generates a label 0 indicating that there is no correspondence between Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_1.

[0038] Next, the training data generation unit 11 determines whether or not there is a correspondence between Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2. That is, it determines whether or not Section 1 of Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_2. As shown in FIG. 7 , if Section 1 of Chapter 1 of the higher-level document 20_1 implies the lower-level document 21_2, the training data generation unit 11 generates a label 1 indicating that Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2 are in a correspondence relationship. If Section 1 of Chapter 1 of the higher-level document 20_1 does not imply the lower-level document 21_2, the training data generation unit 11 generates a label 0 indicating that there is no correspondence between Section 1 of Chapter 1 of the higher-level document 20_1 and the lower-level document 21_2.

[0039] Similarly, the teacher data generation unit 11 determines whether or not there is a correspondence between Section 1 of the upper document 20_1 Chapter 1 and each of the lower documents 21_3 to 21_n. That is, it determines whether or not Section 1 of the upper document 20_1 Chapter 1 implies each of the lower documents 21_3 to 21_n. As shown in FIG. 7 , if Section 1 of the upper document 20_1 Chapter 1 implies each of the lower documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that Section 1 of the upper document 20_1 Chapter 1 and each of the lower documents 21_3 to 21_n are in a correspondence relationship. If Section 1 of the upper document 20_1 Chapter 1 does not imply each of the lower documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that there is no correspondence between Section 1 of the upper document 20_1 Chapter 1 and each of the lower documents 21_3 to 21_n.

[0040] Similarly, the teacher data generation unit 11 determines whether or not there is a correspondence between section 2 of chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. That is, it determines whether or not section 2 of chapter 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n. As shown in FIG. 7 , if section 2 of chapter 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that there is a correspondence between section 2 of chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. If section 2 of chapter 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that there is no correspondence between section 2 of chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n.

[0041] The same applies to each of the sections 3 and onward in chapter 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. The training data generation unit 11 determines whether each of the sections 3 and onward in chapter 1 of the higher-level document 20_1 corresponds to each of the lower-level documents 21_1 to 21_n. If each of the sections 3 and onward in chapter 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 1 for each of the sections. If each of the sections 3 and onward in chapter 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 0 for each of the sections.

[0042] The same applies to each of the sections in the upper document 20_1, Chapter 2 and each of the lower documents 21_1 to 21_n. The training data generation unit 11 determines whether each of the sections in the upper document 20_1, Chapter 2 and each of the lower documents 21_1 to 21_n correspond to each other. If each of the sections in the upper document 20_1, Chapter 2 implies each of the lower documents 21_1 to 21_n, the training data generation unit 11 generates a label 1. If each of the sections in the upper document 20_1, Chapter 2 does not imply each of the lower documents 21_1 to 21_n, the training data generation unit 11 generates a label 0.

[0043] The other sections of the higher-level documents 20_1 to 20_m and the lower-level documents 21_1 to 21_n are similar to those described above, and therefore will not be described again.

[0044] Next, the training data generation unit 11 associates the terms of the higher-level documents 20_1 to 20_m with the lower-level documents 21_1 to 21_n (S13), as shown in Figures 4 and 5. Depending on whether the term in the higher-level document implies the lower-level document, a label is generated indicating whether the term in the higher-level document and the lower-level document correspond to each other. That is, if a sentence (or meaning of the sentence) in the term in the higher-level document implies a sentence (or meaning of the sentence) in the lower-level document, a label is generated indicating that an implication relationship exists between the term in the higher-level document and the lower-level document. Conversely, if a sentence in the term in the higher-level document does not imply a sentence in the lower-level document, a label is generated indicating that an implication relationship does not exist between the term in the higher-level document and the lower-level document.

[0045] Specifically, the training data generation unit 11 determines whether or not Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 corresponds to the lower-level document 21_1. That is, the training data generation unit 11 determines whether or not Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 implies the lower-level document 21_1. As shown in FIG. 8 , if Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 implies the lower-level document 21_1, the training data generation unit 11 generates a label 1 indicating that Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 corresponds to the lower-level document 21_1. If Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 does not imply the lower-level document 21_1, the training data generation unit 11 generates a label 0 indicating that Item 1 of Chapter 1, Section 1 of the higher-level document 20_1 does not correspond to the lower-level document 21_1.

[0046] Next, the training data generation unit 11 determines whether or not Item 1 of the upper document 20_1, Chapter 1, Section 1 corresponds to the lower document 21_2. That is, the training data generation unit 11 determines whether or not Item 1 of the upper document 20_1, Chapter 1, Section 1 implies the lower document 21_2. As shown in FIG. 8 , if Item 1 of the upper document 20_1, Chapter 1, Section 1 implies the lower document 21_2, the training data generation unit 11 generates a label 1 indicating that Item 1 of the upper document 20_1, Chapter 1, Section 1 corresponds to the lower document 21_2. If Item 1 of the upper document 20_1, Chapter 1, Section 1 does not imply the lower document 21_2, the training data generation unit 11 generates a label 0 indicating that Item 1 of the upper document 20_1, Chapter 1, Section 1 does not correspond to the lower document 21_2.

[0047] Similarly, the teacher data generation unit 11 determines whether Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 corresponds to each of the lower-level documents 21_3 to 21_n. That is, the teacher data generation unit 11 determines whether Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 implies each of the lower-level documents 21_3 to 21_n. As shown in FIG. 8 , if Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 implies each of the lower-level documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 corresponds to each of the lower-level documents 21_3 to 21_n. If Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_3 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that Item 1 in Chapter 1, Section 1 of the higher-level document 20_1 does not correspond to each of the lower-level documents 21_3 to 21_n.

[0048] Similarly, the teacher data generation unit 11 determines whether or not there is a correspondence between Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. That is, it determines whether or not Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n. As shown in FIG. 8 , if Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 1 indicating that there is a correspondence between Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. If Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the teacher data generation unit 11 generates a label 0 indicating that there is no correspondence between Item 2 in Chapter 1, Section 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n.

[0049] The same applies to each of the items 3 and onward in Chapter 1, Section 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. The training data generation unit 11 determines whether there is a correspondence between each of the items 3 and onward in Chapter 1, Section 1 of the higher-level document 20_1 and each of the lower-level documents 21_1 to 21_n. If each of the items 3 and onward in Chapter 1, Section 1 of the higher-level document 20_1 implies each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 1 for each of the items. If each of the items 3 and onward in Chapter 1, Section 1 of the higher-level document 20_1 does not imply each of the lower-level documents 21_1 to 21_n, the training data generation unit 11 generates a label 0 for each of the items.

[0050] The other items in the higher-level documents 20_1 to 20_m and the lower-level documents 21_1 to 21_n are similar to those described above, and therefore will not be described again.

[0051] 2.2 Generation Process of Inference Result Data Next, the generation process of the inference result data 16 (inference process) executed by the inference unit 13 will be described with reference to Figures 9 and 10. Figure 9 is a flowchart showing the generation process of the inference result data 16 in the inference unit 13. Figure 10 is a diagram showing the generation process of the inference result data 16. It is assumed that the higher-level document 20a includes a chapter A, which includes a section B, and which includes a requirement item (or clause) C. Below, the process of obtaining requirement item C as a search condition (or search key) and extracting a lower-level document corresponding to requirement item C will be described.

[0052] 9 and 10, the inference unit 13 first acquires chapter A of the higher-level document 20a that includes requirement item C, and extracts lower-level documents 21a that are implied by the acquired chapter A (S21). That is, the inference unit 13 extracts lower-level documents 21a that are associated with chapter A of the higher-level document 20a that includes requirement item C.

[0053] Next, the inference unit 13 obtains a section B of the higher-level document 20a that includes requirement item C. Furthermore, the inference unit 13 extracts a lower-level document 21b that is implied by the obtained section B from among the lower-level documents referenced by the lower-level document 21a extracted in step S21, in other words, from among the lower-level documents that have links set to the lower-level document 21a (S22). That is, the inference unit 13 extracts the lower-level document 21b that is associated with section B of the higher-level document 20a that includes requirement item C.

[0054] Next, the inference unit 13 extracts the lower-level document 21c implied by the requirement item C from among the lower-level documents referenced by the lower-level document 21b extracted in step S22, in other words, from among the lower-level documents linked to the lower-level document 21b (S23). That is, the inference unit 13 extracts the lower-level document 21c associated with the requirement item C. Then, the inference unit 13 outputs inference result data 16 including the extracted lower-level document 21c. This completes the process of generating the inference result data 16 (inference process) by the inference unit 13.

[0055] In the process of generating the inference result data 16 described above, chapter A of the higher-level document 20a containing requirement item C is obtained, the lower-level document 21a implied by the obtained chapter A is extracted, and then section B of the higher-level document 20a containing requirement item C is obtained, and the lower-level document 21a implied by the obtained section B is extracted, but this is not limited to this. It is also possible to first obtain chapter A and section B of the higher-level document 20a containing requirement item C, extract the lower-level document 21a implied by the obtained chapter A, and then extract the lower-level document 21b implied by the obtained section B.

[0056] Next, the generation process (inference process) of the inference result data 16 executed by the inference unit 13 will be described with reference to FIG. 11. FIG. 11 is a diagram showing the generation process of the inference result data 16 from another perspective. Note that the upper-level document 20a includes condition 1, which includes conditions 2 and 3. Condition 2 includes requirement 1, and condition 3 includes requirement 2. Condition 1 corresponds to chapter A of the upper-level document 20a, condition 2 corresponds to section B, and requirement 1 corresponds to requirement item C. Condition 3 corresponds to another section, and requirement 2 corresponds to another requirement item. Below, a process for extracting a lower-level document corresponding to requirement 1 from condition 1 of the upper-level document 20a that includes requirement 1 as a search condition will be described.

[0057] 11, the inference unit 13 first acquires condition 1 of the higher-level document 20a that includes requirement 1, and extracts the lower-level document 21a that is implied by the acquired condition 1. That is, the inference unit 13 extracts the lower-level document 21a that is associated with condition 1 of the higher-level document 20a that includes requirement 1.

[0058] Next, the inference unit 13 acquires condition 2 of the higher-level document 20a, which includes requirement 1. Furthermore, the inference unit 13 extracts lower-level documents 21b implied by the acquired condition 2 from among the lower-level documents referenced by the acquired lower-level document 21a, in other words, from among the lower-level documents linked to the lower-level document 21a. In other words, the inference unit 13 extracts lower-level documents 21b associated with condition 2 of the higher-level document 20a, which includes requirement 1.

[0059] Next, the inference unit 13 extracts the lower-level document 21c implied by requirement 1 from among the lower-level documents referenced by the extracted lower-level document 21b, in other words, from among the lower-level documents linked to the lower-level document 21b. That is, the inference unit 13 extracts the lower-level document 21c associated with requirement 1. The inference unit 13 then outputs inference result data 16 including the extracted lower-level document 21c. This completes the process of generating the inference result data 16 (inference process) by the inference unit 13.

[0060] As described above, in the embodiment, lower-level documents are searched in the order of chapter, section, and paragraph of the upper-level document in which requirement item C (or requirement 1) is described, and lower-level documents that meet requirement item C are extracted and output.

[0061] 3. Hardware Configuration of the Embodiment Next, the hardware configuration of the natural language processing apparatus 10 of the embodiment will be described with reference to Fig. 12. Here, an example in which the natural language processing apparatus 10 is configured by a computer 30 will be described.

[0062] 12 is a diagram showing an example of the hardware configuration of the natural language processing apparatus 10. The natural language processing apparatus 10 (i.e., a computer 30) has a processor 31, a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, an auxiliary storage device 34, and an input / output interface 35. The processor 31, the ROM 32, the RAM 33, the auxiliary storage device 34, and the input / output interface 35 are electrically connected to one another via a bus 36, and are capable of exchanging data and signals with one another via the bus 36.

[0063] The processor 31 is configured by a general-purpose hardware processor including, for example, a CPU (Central Processing Unit), a GPU (Graphical Processing Unit), etc. The processor 31 controls the ROM 32, the RAM 33, the auxiliary storage device 34, and the input / output interface 35 as a whole.

[0064] The ROM 32 is a non-volatile memory that constitutes part of the main storage device. The ROM 32 non-temporarily stores a startup program required when starting up the natural language processing device 10. The natural language processing device 10 starts up when the processor 31 executes the program in the ROM 32. The ROM 32 is configured, for example, by an EPROM (Erasable Programmable Read Only Memory), and stores various settings at startup in addition to the startup program.

[0065] The RAM 33 is a volatile memory that constitutes part of the main storage device. The RAM 33 temporarily stores programs required for processing by the processor 31 and data required for executing the programs. The processor 31 executes the programs in the RAM 33 to perform calculations on the data in the RAM 33 and store the calculation results in the RAM 33.

[0066] The auxiliary storage device 34 is configured by a non-volatile memory such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The auxiliary storage device 34 non-temporarily stores programs to be executed by the processor 31 and data required for executing the programs. The processor 31 loads the programs and data in the auxiliary storage device 34 into the RAM 33 and executes the programs to perform various functions.

[0067] The input / output interface 35 is connected to an external input device 41, an output device 42, etc., and enables input of information from the input device 41 and output of information to the output device 42. For example, the input / output interface 35 may be a wired interface or a wireless interface. A wired interface includes a port to which a device is connected. A wireless interface includes Bluetooth (registered trademark), Wi-Fi (registered trademark), etc. The input device 41 may include a keyboard, a mouse, a touch panel, a receiving device, a disk drive, etc. The input device 41 is not limited to these, and may include any other input device. The output device 42 may include a display, a transmitting device, a disk drive, etc. The output device 42 is not limited to these, and may include any other output device. The input device 41 and the output device 42 may be configured as an input / output device 43 that has the functions of both the input device 41 and the output device 42.

[0068] Input data, such as data relating to higher-level documents and lower-level documents, and data such as requirement items of the higher-level documents to be searched, is input to the teacher data generation unit 11 or the inference unit 13 via the input device 41 .

[0069] The program non-temporarily stored in the auxiliary storage device 34 is provided to the computer 30, for example, via a recording medium 44 on which the program is non-temporarily recorded and which can be read by the computer 30. Such a recording medium 44 is called a non-temporarily computer-readable recording medium. Non-temporarily computer-readable recording media include disks such as flexible disks, optical disks (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), and magneto-optical disks (MO, etc.), as well as semiconductor memories.

[0070] If the recording medium 44 is a disk, the program non-temporarily stored in the auxiliary storage device 34 is read into the auxiliary storage device 34 via the disk drive serving as the input device 41 and the input / output interface 35, or if the recording medium 44 is a semiconductor memory, via a port serving as the input / output interface 35, and non-temporarily stored therein. Alternatively, the program may be stored in a server on a network, downloaded from the server, and non-temporarily stored in the auxiliary storage device 34.

[0071] When the computer 30 starts up, the processor 31 executes a program in the ROM 32 and loads and starts the OS into the RAM 33. Under control of the OS, the processor 31 monitors input instructions, connections to external devices, and the like. Under control of the OS, the processor 31 also sets up a program area and a data area in the RAM 33. In response to an input instruction to start the natural language processing device 10, the processor 31 loads a natural language processing program from the auxiliary storage device 34 into the program area of ​​the RAM 33 and loads data required for executing the natural language processing program from the auxiliary storage device 34 into the data area of ​​the RAM 33. The processor 31 calculates data in the data area according to the natural language processing program and writes the calculation results to the data area. Through these operations, the processor 31, the RAM 33, the auxiliary storage device 34, the input / output interface 35, and the bus 36 work together to execute at least some of the functions of the components of the natural language processing device 10, namely, the teacher data generation unit 11, the pre-learning generation unit 12, and the inference unit 13.

[0072] 4. Effects of the Embodiments, etc. According to the embodiments of the present invention, it is possible to provide a natural language processing device that can search for lower-level documents that satisfy requirements taking into account the context of the chapters, sections, and paragraphs of the higher-level document in which the requirements are written.

[0073] In the embodiment, the training data generation unit 11 generates labels indicating the correspondence between chapters of a higher-level document and lower-level documents, labels indicating the correspondence between sections of the higher-level document and lower-level documents, and labels indicating the correspondence between terms of the higher-level document and lower-level documents, thereby generating training data 14. Next, the pre-learning generation unit 12 uses the training data 14 as input data and a machine learning model (e.g., BERT) to generate an entailment learning model 15 for determining the entailment relationship between each of the chapters, sections, and terms of the higher-level document and the lower-level documents. The inference unit 13 then acquires information about the chapters and sections of the higher-level document, including the requirement items as search criteria, and, based on the entailment learning model 15 generated using the training data 14, searches for and narrows down the lower-level documents in the order of chapters, sections, and requirement items of the higher-level document according to the reference relationships between the lower-level documents. This allows lower-level documents corresponding to the requirement items, taking into account the context of the chapters, sections, and requirement items of the higher-level document, including the requirement items.

[0074] In other words, in the configuration of this embodiment, training data 14 is created to which labels indicating the correspondence between each chapter, section, and paragraph of a higher-level document and a lower-level document are assigned, and an implication learning model 15 is generated that infers the correspondence between each chapter, section, and paragraph of the higher-level document and a lower-level document. During inference, information on the chapter and section containing the requirement items of the higher-level document to be searched is obtained, and the lower-level documents are searched and narrowed down in the order of chapter, section, and requirement item while taking into account the reference relationships between the lower-level documents. This makes it possible to search for lower-level documents that satisfy the requirement items while taking into account the context of the chapter, section, and requirement item.

[0075] The functional blocks described in the above embodiments can be realized as either hardware or computer software, or a combination of both. It is not necessary for the functional blocks to be distinguished as in the above examples. For example, some functions may be performed by functional blocks other than the illustrated functional blocks. Furthermore, the illustrated functional blocks may be further divided into smaller functional sub-blocks. Furthermore, the order of processing in the flowcharts described in the above embodiments can be changed as much as possible.

[0076] It should be noted that the present invention is not limited to the above-described embodiment, and can be practiced in various modified forms without departing from the spirit and scope of the present invention.

[0077] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

[0078] 10...Natural language processing device 11...Teacher data generation unit 12...Pre-learning generation unit 13...Inference unit 14...Teacher data 15...Entailment relationship learning model 16...Inference result data 20...Superior document 20a...Superior document 21...Subordinate document 21a...Subordinate document 21b...Subordinate document 21c...Subordinate document 30...Computer 31...Processor 32...ROM 33...RAM 34...Auxiliary storage device 35...Input / output interface 36...Bus 41...Input device 42...Output device 43...Input / output device 44...Recording medium.

Claims

1. A natural language processing device comprising: a teacher data generation unit that generates teacher data representing the correspondence between each of the chapters, sections, and items included in the upper document and the lower document; and an inference unit that acquires the chapters and sections of the first upper document including the item as a search condition, and performs a search for the lower document in the order of the chapter, the section, and the item of the first upper document according to the reference relationship between the lower documents based on a machine learning model generated using the teacher data, and outputs the lower document corresponding to the item of the first upper document.

2. The natural language processing device according to claim 1, wherein the teacher data generation unit generates, as the teacher data, a matrix representing an implicative relationship between the chapter of the upper document and the lower document, a matrix representing an implicative relationship between the section of the upper document and the lower document, and a matrix representing an implicative relationship between the item of the upper document and the lower document.

3. The natural language processing device according to claim 1, wherein the inference unit extracts a first lower document implied by the chapter of the first upper document from among the lower documents, extracts a second lower document implied by the section of the first upper document from among the lower documents referred to from the first lower document, extracts a third lower document implied by the item of the first upper document from among the lower documents referred to from the second lower document, and outputs the extracted third lower document.

4. The natural language processing device according to claim 1, further comprising a pre-learning generation unit that generates an implicative relationship learning model for determining an implicative relationship between each of the chapters, sections, and items of the upper document and the lower document using the teacher data as the machine learning model.

Citation Information

Patent Citations

  • Method of determining document similarity

    JP2015219799A

  • Document management device and document management method

    WO2017149711A1