Document analysis system, document analysis method, and program

JP7927222B2Active Publication Date: 2026-10-01KYUSHU UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022152232
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2026-10-01
Estimated Expiration
2042-09-26

AI Technical Summary

Benefits of technology

【0019】 本発明により、技術文書の理解を容易にする文書解析システム等が提供される。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927222000001
    Figure 0007927222000001
  • Figure 0007927222000002
    Figure 0007927222000002
  • Figure 0007927222000003
    Figure 0007927222000003
Patent Text Reader

Abstract

To provide a document analysis system and the like to facilitate understanding of a technical document.SOLUTION: A document analysis system 1 includes an annotation unit 16 that accepts annotation setting of a unique expression tag for a limited document group which is a document group of a specific technical field, generates annotated annotation data 5, and can set a delimiter tag S given to character information that delimits a unique expression tag string forming a certain one semantic relationship as a unique expression tag, a unique expression extraction model learning unit 17 that generates a unique expression extraction model 8 using the annotation data 5, a document input unit 11 that inputs a document of the technical field as an analysis target, and a unique expression extraction unit 12 that extracts information on the delimiter tag S along with information on another unique expression tag from the document of the analysis target using the unique expression extraction model 8.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a document analysis system for technical documents, etc. [Background Art]

[0002] Extracting required technical elements from a huge volume of technical documents is a major challenge. For example, technical documents include many intellectual property-related documents, starting with patent publications. These intellectual property-related documents have been increasingly digitized and further classified, and machine search based on such classifications and free keywords has made it easier to extract a target document group.

[0003] In recent years, by performing machine learning on documents, an analysis method that organizes and outputs required information has been developed. For example, BERT (Bidirectional Encoder Representations from Transformers) has drawn attention as a development of technology, which can learn various arrangement positions of words used in sentences and similar expressions, and assign appropriate named entity tags to words and phrases (Non-Patent Document 1). In addition, there is also a report on PatentBERT focusing on patent data as intellectual property-related documents (Non-Patent Document 2). [Prior Art Documents] [Non-Patent Documents]

[0004] [Non-Patent Document 1] Gururangan, S., Marasovic, A., Swayamdipta, S., Lo, K., Beltagy,I., Downey, D., & Smith, NA (2020, July). Don'tStop Pretraining: Adapt Language Models to Domains and Tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp.8342-8360). [Non-Patent Document 2] Lee, J.-S. and Hsiang, J.: Patentbert: Patent classificationwithfine-tuning a pre-trained bert model, arXiv preprint arXiv:1906.02124 (2019). [Overview of the project] [Problems that the invention aims to solve]

[0005] The analysis method described above can extract appropriate named entities, thus reducing the likelihood of missed results due to simple keyword searches. Furthermore, it can automate the manual classification process for intellectual property-related documents. In this way, recent machine learning methods seem to be a panacea for analyzing technical documents. However, technical documents often employ specific expression methods depending on the field, leaving a challenge for further improvement in analysis accuracy.

[0006] Furthermore, for researchers, simply extracting relevant documents is not enough to achieve their objectives; they must read and understand the documents themselves to grasp what technologies are disclosed. Additionally, those seeking to verify intellectual property rights need to quickly understand the scope of rights disclosed in intellectual property-related documents such as patent documents in order to prevent patent infringement and to formulate new research areas.

[0007] This invention has been made in view of the aforementioned problems, and aims to provide a document analysis system and the like that facilitates the understanding of technical documents. [Means for solving the problem]

[0008] The first invention for achieving the aforementioned objective is an annotation unit that accepts annotation settings for named entity tags for a limited group of documents, which is a group of documents in a specific technical field, and generates annotated training data, wherein the annotation unit is capable of setting delimiter tags to be attached to character information that separates a sequence of named entity tags that form a certain semantic relationship as the named entity tags, The aforementioned delimiter tag is included as a named entity tag. The document analysis system is characterized by comprising: a named entity recognition model learning unit that generates a named entity recognition model using the aforementioned training data; a document input unit that inputs documents in the aforementioned technical field as the target for analysis; and a named entity recognition unit that uses the named entity recognition model to extract information of delimiter tags together with information of other named entity tags from the target documents.

[0009] According to the first invention, the system accepts annotation settings for named entity tags for a limited set of documents, which are a group of documents in a specific technical field, and generates a named entity recognition model using annotated training data. In particular, in this invention, when setting annotations, it is possible to set delimiter tags that are attached to character information that separates a sequence of named entity tags that form a certain semantic relationship, and the named entity recognition model extracts information of the delimiter tags along with information of other named entity tags from the document to be analyzed. Since semantic units (examples of named entity tags) can be identified from the document based on the extracted delimiter tags, understanding and analyzing technical documents becomes easier.

[0010] Furthermore, by referring to the delimiter tag extracted by the named entity recognition unit, Each sequence of named entity tags, separated by delimiter tags, is obtained as a sequence of named entity tags forming a single semantic relationship. The system may further include a structural analysis unit that performs structural analysis of the aforementioned document. A sequence of named entity tags with semantic relationships can be clearly identified. This allows for easy structural analysis of the document being analyzed.

[0011] Furthermore, the annotation unit accepts annotation settings for named entity tags that have a master-slave relationship. , the named entity tags having the aforementioned master-slave relationship are included in the training data and are subject to learning and extraction for named entity recognition. This is also acceptable. This allows for annotation settings that establish a master-slave relationship between named entity tags. Furthermore, it enables named entity recognition that reflects the master-servant relationship. can.

[0012] Furthermore, the annotation unit accepts annotation settings for related named entity tags for cue words that associate named entity tags, The aforementioned related named entity tags may be included in the training data and used for learning and extraction of named entity recognition. Associate named entity tags with each other. clue For a word Named entities Set tags The learning targets for named entity recognition and This improves the performance of extracting named entity tags.

[0013] Also, The structural analysis unit refers to a tag column pattern table that stores pre-associated patterns of named entity tag columns with data formats of structured data, identifies the pattern of each named entity tag column by pattern matching, and converts the information of each named entity tag column into structured data in the corresponding data format. You may do so. This allows for the efficient conversion of named entity tag sequences with semantic relationships into structured data, enabling efficient structural analysis of documents.

[0014] Also, The aforementioned tag column pattern table holds multiple patterns with different word orders for named entity tag columns that represent the same meaning. You may This allows for flexible handling of expressions with the same meaning but different word orders.

[0015] Also, The aforementioned named entity tags include element tags, numeric tags, and unit tags. That's fine. This allows for the effective extraction and structuring of combinations of elements, numerical values, and units that frequently appear in technical documents in fields such as chemistry.

[0016] Furthermore, it is desirable that the aforementioned limited document group and the document to be analyzed are intellectual property-related documents. Intellectual property-related documents are generally complex, but by applying the named entity recognition model trained with the delimiter tags of the present invention, understanding the documents becomes easier.

[0017] The second invention is an annotation process in which a computer accepts annotation settings for named entity tags for a limited set of documents which is a set of documents in a specific technical field, and generates annotated training data, wherein the annotation process is capable of setting delimiter tags to be attached to character information that separates a sequence of named entity tags which forms a certain semantic relationship as the named entity tags, The aforementioned delimiter tag is included as a named entity tag.a named entity extraction model learning step of generating a named entity extraction model using the teacher data; a document input step of inputting a document in said technical field as an analysis target; and a named entity extraction step of extracting information on delimiter tags from said document to be analyzed together with information on other named entity tags using said named entity extraction model; wherein the document analysis method is characterized by executing the steps.

[0018] The third invention provides a program that causes a computer to function as: an annotation unit that accepts annotation settings for named entity tags for a limited document group, which is a document group of a specific technical field, and generates annotated teacher data, wherein said annotation unit is capable of setting a delimiter tag to be added to character information that delimits a named entity tag sequence forming a certain semantic relationship as said named entity tag; The aforementioned delimiter tag is included as a named entity tag. a named entity extraction model learning unit that generates a named entity extraction model using said teacher data; a document input unit that inputs a document in said technical field as an analysis target; and a named entity extraction unit that extracts information on delimiter tags from said document to be analyzed together with information on other named entity tags using said named entity extraction model. [Effects of the Invention]

[0019] The present invention provides a document analysis system and the like that facilitates understanding of technical documents. [Brief Description of the Drawings]

[0020] [Figure 1] It is a diagram showing the learning function of the document analysis system 1. [Figure 2] It is a diagram showing an example of annotation settings for each named entity tag including a delimiter tag. [Figure 3] It is a diagram showing an example of annotation settings for a main named entity tag, a subordinate named entity tag, and a related named entity tag. [Figure 4] It is a diagram showing the re-learning function of the document analysis system 1. [Figure 5] It is a diagram showing the analysis function of the document analysis system 1. [Figure 6]This figure shows an example of named entity recognition. [Figure 7] This figure shows an example of the tag column pattern table 90. [Figure 8] This figure shows an example of structural analysis. [Figure 9] This figure shows an example of the data structure of document extraction DB100. [Figure 10] This figure shows an example of named entity recognition when no delimiter tags are set. [Figure 11] This figure shows an example of structural analysis when delimiter tags are not set. (a) shows a correct structural analysis example, and (b) shows an incorrect structural analysis example. [Figure 12] This is a diagram showing the hardware configuration of computer 30. [Figure 13] This is a flowchart showing the learning process flow. [Figure 14] This is a flowchart showing the flow of the retraining process. [Figure 15] This is a flowchart showing the flow of the analysis process. [Figure 16] This diagram shows the search function of document analysis system 1. [Figure 17] This is a flowchart showing the search process flow. [Modes for carrying out the invention]

[0021] Preferred embodiments of the present invention will be described in detail below with reference to the drawings.

[0022] (1. Learning function) First, referring to Figure 1, we will explain the learning function for training the named entity recognition model 8 used for document analysis. The learning function consists of "pre-training," "re-pre-training," and "actual training," and the named entity recognition model 8 is generated by performing these training processes. Each learning function will be explained below.

[0023] (1-1. Pre-learning) Pre-training is performed by the pre-training unit 14 of the document analysis system 1. Specifically, as shown in Figure 1, the pre-training unit 14 performs unsupervised pre-training using a large set of documents 3 to generate a pre-trained model 6.

[0024] Large-scale document collection 3 is a database (a large-scale corpus for general purposes) that has been structured and compiled on a large scale from texts written on the web, in books, etc., and serves as training data for training the pre-trained model 6. In this embodiment, a database (a large-scale corpus) that has been structured and compiled on a large scale from articles on the Japanese Wikipedia is used as the large-scale document collection 3.

[0025] The training of the pre-trained model 6 performed by the pre-training unit 14 is preferably based on neural network-based deep learning. In this case, using a Transformer as the deep learning model enables parallel processing and shortens the training time. Such deep learning can be performed using a known natural language processing technique, such as BERT. In this embodiment, the pre-training unit 14 uses BERT technology to train the pre-trained model 6 using a large amount of document data obtained from Japanese Wikipedia.

[0026] Furthermore, since pre-training is a computationally intensive process using a large amount of data, it is desirable that it be performed in advance. That is, it is desirable that the pre-trained model 6 be prepared in advance and stored in the storage device 32 (Figure 12) of the computer 30. In this case, the function and process for training the pre-trained model 6 can be omitted from this embodiment.

[0027] (1-2. Re-learning) Re-pre-training is performed by the re-pre-training unit 15 of the document analysis system 1. Specifically, as shown in Figure 1, the re-pre-training unit 15 re-pre-trains the pre-trained model 6 using a limited document group 4A, which is a set of documents in a specific technical field, to generate a re-pre-trained model 7. Here, documents in a technical field include, for example, technical papers, technical reports, and intellectual property-related documents.

[0028] The restricted document group 4A is a database (special-purpose corpus) that structures and collects documents in a specific technical field to be analyzed, and serves as training data for re-pre-training the pre-trained model 6. The restricted document group 4A consists of documents in the same technical field as the target document 2 (see Figure 5) to be analyzed.

[0029] Furthermore, it is desirable that the restricted document group 4A consists of intellectual property-related documents. Intellectual property-related documents are documents that contain the contents of patent applications or utility model registration applications, such as published patent gazettes, patent gazettes (patent publication gazettes), and registered utility model gazettes. If the restricted document group 4A consists of intellectual property-related documents, the specific technical field is, for example, a technical field classified by patent classification symbols consisting of IPC, FI, F-terms, etc.

[0030] Re-pre-training can be performed using a known technique known as BERT, similar to pre-training. By re-pre-training the pre-trained model 6 using the limited document set 4A, a learning model (re-pre-trained model 7) adapted to the analysis of a specific technical field can be obtained. Furthermore, the use of the limited document set 4A efficiently improves the learning accuracy achieved through large-scale pre-training.

[0031] While it is possible to train a pre-trained model using BERT from the beginning with a limited set of documents 4A instead of the large set of documents 3, this would require preparing and working with a large number of training samples within a limited scope, which is inefficient. In this respect, improving accuracy through further re-pre-training based on a generally usable pre-trained model 6, as in this embodiment, is effective in achieving both accuracy and efficiency.

[0032] Furthermore, re-pre-training is not required. That is, the pre-trained model 6 may be prepared in advance and stored in the storage device 32 (Figure 12) of the computer 30. In this case, the function and process for training the re-pre-trained model 7 can be omitted from this embodiment.

[0033] (1-3. Practical Learning) The actual training consists of a process to generate training data (annotation data) for training the named entity recognition model 8, and a process to fine-tune the pre-trained model 6 or the re-pre-trained model 7 using the training data (annotation data) to generate the named entity recognition model 8. Specifically, the actual training is performed by the annotation unit 16 and the named entity recognition model training unit 17 of the document analysis system 1.

[0034] The annotation unit 16 receives annotation settings for various named entity tags from the user for the limited document group 4B, which is a group of documents in a specific technical field, and generates annotated training data (hereinafter referred to as "annotation data 5").

[0035] Restricted document group 4B consists of documents in the same technical field as the subject document 2 and restricted document group 4A, which are the subjects of analysis. Furthermore, it is desirable that restricted document group 4B consists of intellectual property-related documents. Additionally, the documents in restricted document group 4B may be identical to, or partially identical to, the documents in restricted document group 4A.

[0036] In this embodiment, the annotation unit 16 can set delimiter tags as named entity tags, which are attached to character information (characters, symbols, or combinations thereof) that separates a sequence of named entity tags that form a certain semantic relationship.

[0037] Figure 2 illustrates an example of annotation settings for each named entity tag, including delimiter tags. Figure 2 shows an example of annotation settings for the following claim example from the "Claims" section of an intellectual property document: "A steel material having a composition consisting of 0.03-1.5% by mass of C, 5-10% of Ni or 0.7-2% of Cu, and further comprising 0.1-4% of Co and 5% or less of Si."

[0038] In the example in Figure 2, the user has set five types of named entity tags for a sentence. "E" is an "element tag" set for element names, "LF" is a "lower limit tag" set for the lower limit of a numerical range, "UF" is an "upper limit tag" set for the upper limit of a numerical range, and "U" is a "unit tag" set for the numerical unit representation. "S" represents a "delimiter tag".

[0039] The delimiter tag S is set as character information that separates a sequence of named entity tags that form a certain semantic relationship. In the case of Figure 2, (a) "C at 0.03-1.5%" (Named entity tag sequence: "E", "LF", "UF", "U") (b) "Ni 5-10%" (Named entity tag sequence: "E", "LF", "UF", "U") (c) "Cu at 0.7-2%" (Named entity tag sequence: "E", "LF", "UF", "U") (d) "Co at 0.1-4%" (Named entity tag sequence: "E", "LF", "UF", "U") (e) 'Si content of 5% or less' (Named entity tag sequence: "E", "UF", "U") This is a description that specifies the content of each element contained in the steel material, and (a) to (e) above each form a single semantic relationship. The delimiter tag S is set as the character information that separates each of the named entity tag sequences in (a) to (e) above.

[0040] For example, in Figure 2, delimiter tags S are set for the character information that separates each of the named entity tag sequences (a) to (e) above: 'and', ',' 'or', 'including', 'furthermore', ',' and 'consisting of'. Note that the character information for which delimiter tags S are set is not limited to specific characters or symbols, but may be any characters or symbols that separate a named entity tag sequence that forms a single semantic relationship.

[0041] The annotation unit 16 may also accept annotation settings for named entity tags that have a master-slave relationship. Specifically, the annotation unit 16 accepts annotation settings that associate a master named entity tag T1 with a subordinate named entity tag T2 that is subordinate to the master named entity tag T1.

[0042] The primary named entity tag T1 is a named entity tag set for a named entity in a document within the restricted document group 4B. The secondary named entity tag T2 is one or more named entity tags set for a named entity in a document within the restricted document group 4B, subordinate to the primary named entity tag T1. The set primary named entity tag T1 and secondary named entity tags T2 are associated with each other and stored in the annotation data 5.

[0043] Specifically, if the limited document group 4B is an intellectual property-related document, such as the "claims," ​​a primary named entity tag T1 can be set for named entities relating to specific inventive features described in the claims, and a secondary named entity tag T2 can be set for named entities that more specifically represent the nature, structure, etc., of the inventive features to which the primary named entity tag T1 is set.

[0044] For example, in the field of chemistry, a primary named entity tag T1 is set for named entities (inventive features) that represent specific elements or properties, and a secondary named entity tag T2 is set in association with the primary named entity tag T1 for named entities that represent the scope or degree of specific elements or properties. This allows information about the scope and degree (secondary named entities) disclosed in relation to specific elements or properties (primary named entities) to be learned, enabling a detailed analysis of the target document 2. This type of tagging is particularly effective in intellectual property-related documents, as the scope of rights is often expressed in terms of primary and secondary relationships.

[0045] The annotation unit 16 may also accept the setting of related named entity tags Tr from the user for cue words that associate named entity tags. For example, suppose a primary named entity tag T1 is set for named entities related to a specific element, and a secondary named entity tag T2 is set for named entities related to the numerical range of that specific element. In this case, related named entity tags Tr are set for named entities (cue words) that supplement the range or degree, such as symbols or phrases that indicate the upper and lower limits of the numerical range. By setting related named entity tags Tr in this way, further improvement in the performance of the named entity extraction model 8 can be expected.

[0046] Figure 3 illustrates an example of annotation settings for the primary named entity tag T1, secondary named entity tag T2, and related named entity tag Tr. In Figure 3, an example is shown in which the same claim example as in Figure 2 is annotated with the element name as the primary named entity and the lower limit value, upper limit value, and unit expression as secondary named entities.

[0047] Specifically, the element name is set to the element tag T1_E, which is the primary named entity tag T1. The lower limit value tag T2_LF, upper limit value tag T2_UF, and unit tag T2_U are set to the lower limit value tag T2_LF, upper limit value tag T2_UF, and unit tag T2_U, respectively, which are secondary named entity tags T2. In addition, the range tag Tr_R, which is the associated named entity tag Tr, is set as a cue word to associate the lower limit value tag T2_LF and the upper limit value tag T2_UF with the range symbol "~".

[0048] The named entity recognition model learning unit 17 generates a named entity recognition model 8 by fine-tuning the pre-trained model 6 or the re-pre-trained model 7 with annotation data 5 (training data in which named entity tags are set for the limited document group 4B). Fine-tuning can be performed using a known technique known as BERT.

[0049] In this embodiment, a language model (pre-trained model 6) is pre-trained using unsupervised data from Japanese Wikipedia, and then a language model (re-pre-trained model 7) is re-pre-trained using the restricted document set 4A. After that, the parameters of the re-pre-trained model 7 are adjusted by fine-tuning using annotation data 5 (training data), which is the dataset for the named entity recognition task, to generate a named entity recognition model 8.

[0050] The learning function described above generates a named entity recognition model 8 that extracts named entity information from documents. Alternatively, the named entity recognition model 8 may be generated without using the large document set 3, the pre-trained model 6, or the re-pre-trained model 7. That is, the named entity recognition model 8 may be directly trained and generated using machine learning based on annotation data 5, which is created by applying annotation settings to the limited document set 4B.

[0051] (1-4. Re-learning) Next, with reference to Figure 4, we will describe the re-training process to further improve the learning accuracy of the named entity recognition model 8. As shown in Figure 4, the re-training is performed by the document input unit 11, named entity recognition unit 12, re-annotation unit 18, and named entity recognition model re-fine-tuning unit 19, which are provided by the document analysis system 1.

[0052] The document input unit 11 accepts input of a limited document group 4C, which is a set of documents in a specific technical field. The limited document group 4C is a set of documents in the same technical field as target document 2 and limited document groups 4A and 4B. Furthermore, it is desirable that the limited document group 4C consists of intellectual property-related documents.

[0053] The named entity recognition unit 12 extracts named entity information from each document of the restricted document group 4C using the named entity recognition model 8 and outputs extracted data 9. In this embodiment, the extracted named entity information includes information for each named entity tag, including the delimiter tag S.

[0054] The re-annotation unit 18 displays the extracted data 9 on the display device 34, accepts tag correction operations (correction operations for each named entity tag including the delimiter tag S) for each extracted named entity tag, and regenerates the limited document group 4C (hereinafter referred to as "re-annotation data 20") with the corrected named entity tags set as training data.

[0055] The named entity recognition model refining unit 19 refines the named entity recognition model 8 using the re-annotated data 20. Refining can be performed using a known technique known as BERT. Refining can be performed repeatedly until a predetermined learning performance is obtained.

[0056] Furthermore, while a model suitable for analyzing target document 2 is constructed through re-pre-training, actual training, and re-actual training, overfitting may reduce the accuracy of extraction from target document 2. Therefore, it is desirable to appropriately perform generalization performance evaluation through cross-validation at each training stage of re-pre-training, actual training, and re-actual training.

[0057] (2. Analysis function) Next, with reference to Figure 5, the analysis functions of the document analysis system 1 will be described. The analysis functions mainly consist of a document input unit 11, a named entity recognition unit 12, and a structural analysis unit 13.

[0058] The document input unit 11 accepts input of one or more target documents 2, which are the documents to be analyzed. The target documents 2 are documents in a specific technical field to be analyzed, and are in the same technical field as the limited document groups 4A, 4B, and 4C. It is also desirable that the target documents 2 are intellectual property-related documents.

[0059] The named entity recognition unit 12 extracts named entity information from the target document 2 using the named entity recognition model 8, and outputs the extracted data 10 containing the named entity information to the structural analysis unit 13. In this embodiment, the extracted named entity information includes information for each named entity tag, including the delimiter tag S.

[0060] Figure 6 shows an example of named entity recognition. Figure 6 shows an example of named entity recognition performed on the example claim sentence in the "Claims" section of an intellectual property-related document: "A steel material having a composition consisting of 0.01-2.0% C, 3-10% Ni or 0.5-1% Cu by mass, and further comprising 0.1-3% Co and 3% or less Si." As shown in the figure, the named entity recognition result yields extracted data 10 in which information for each named entity tag, including the delimiter tag S, is extracted.

[0061] The structural analysis unit 13 performs structural analysis of the target document 2 based on the extracted data 10 extracted by the named entity recognition unit 12 and generates structured data 40. Specifically, the structural analysis unit 13 obtains each named entity tag sequence separated by the delimiter tag S from the delimiter tag S contained in the extracted data 10 as a named entity tag sequence (meaningful unit) that forms a single semantic relationship. Then, the structural analysis unit 13 refers to the tag sequence pattern table 90 (Figure 7), identifies the pattern of each obtained named entity tag sequence by pattern matching, and converts the information of each named entity tag sequence into structured data 40 in the data format corresponding to each pattern.

[0062] Figure 7 shows an example of the tag column pattern table 90. The tag column pattern table 90 is a table that stores pre-associated patterns of named entity tag columns with data formats for structured data. As shown in Figure 7, the data format for structured data is predetermined for each pattern of named entity tag columns. For example, pattern 1 is a named entity tag column pattern consisting of "element tag E" → "lower limit numerical tag LF" → "upper limit numerical tag UF" → "unit tag U", and the data format for structured data of this pattern is (E: LF, UF, U).

[0063] Furthermore, when specifying the element content in a claim, it may be written in the order of element → numerical value, such as "contains 0.3% to 1% Cu," or in the order of numerical value → element, such as "contains 0.3% to 1% Cu." In either case, it is necessary to ensure that the content can be recognized by pattern matching.

[0064] Therefore, the tag column pattern table 90 shown in Figure 7 includes patterns where element tag E appears at the beginning of the tag column, such as pattern 1 ("element tag E → lower limit numerical tag LF → upper limit numerical tag UF → unit tag U") and pattern 3 ("element tag E → upper limit numerical tag UF → unit tag U"), and patterns where element tag E appears at the end of the tag column, such as pattern 2 ("lower limit numerical tag LF → upper limit numerical tag UF → unit tag U → element tag E") and pattern 4 ("upper limit numerical tag UF → unit tag U → element tag E").

[0065] Furthermore, the patterns in the tag column pattern table 90 are not limited to the example in Figure 7, and any named entity tag column patterns can be stored in association with the data format of the structured data.

[0066] Figure 8 shows an example of document structure analysis. Figure 8 shows an example of structure analysis performed based on the extracted data 10 in Figure 6. As shown in Figure 8, the structure analysis unit 13 first obtains each named entity tag sequence Tc1 to Tc5, separated by delimiter tags S, from the named entity tag sequence contained in the extracted data 10, as a named entity tag sequence (meaningful unit) that forms a single semantic relationship. Then, the structure analysis unit 13 refers to the tag sequence pattern table 90 (Figure 7) to identify the pattern of each named entity tag sequence Tc1 to Tc5 by pattern matching, and converts the information of each named entity tag sequence Tc1 to Tc5 into structured data 40 in a data format corresponding to each pattern.

[0067] For example, named entity tag columns Tc1~Tc4 are tag columns consisting of "element tag E" → "lower limit numerical tag LF" → "upper limit numerical tag UF" → "unit tag U", and correspond to pattern 1 of tag column pattern table 90 (Figure 7). Therefore, the information of named entity tag columns Tc1~Tc4 is processed according to the data format (E: LF, UF, U) corresponding to pattern 1, (C: This is converted into structured data 40 as follows: 0.01, 2.0, mass%), (Ni: 3, 10, mass%), (Cu: 0.5, 1, mass%), (Co: 0.1, 3, mass%).

[0068] Furthermore, the named entity tag column Tc5 is a tag column consisting of "element tag E" → "upper value tag UF" → "unit tag U", and corresponds to pattern 3 of the tag column pattern table 90. Therefore, the information of the named entity tag column Tc5 is in accordance with the data format (E: , UF, U) corresponding to pattern 3, (Si: This is converted into structured data 40, which is , 3, mass%).

[0069] The structural analysis unit 13 stores the generated structured data 40 in the document extraction DB 100, linking it with a document ID and document attributes that uniquely identify the document.

[0070] Figure 9 shows an example of the data structure of the document extraction DB 100 that stores structured data 40. As shown in Figure 9, structured data 40 is stored for each document, linked to the document ID 41, document name 43, document attributes 45, etc. For example, in the case of a document with document ID 41 "T001" (document name 43: claims, document attributes 45: patent publication), the structured data 40 stored is (C: 0.01, 2.0, mass%), (Cu: 0.5, 1, mass%), (Co: 0.1, 3, mass%), (Si: , 3, mass%)...

[0071] For comparison, we will now supplement the explanation of document structure analysis when named entity recognition is performed without setting a delimiter tag S, unlike in this embodiment. Figure 10 shows an example of extracted data 10 from which named entity information has been extracted. The example sentences to be extracted are the same as in Figure 6. As shown in Figure 10, since the delimiter tag S is not set, the delimiter tag S is naturally not extracted. Named entity tags other than the delimiter tag S are extracted in the same way as the extraction results in Figure 6.

[0072] Figure 11 shows an example of performing structural analysis on the extracted data 10 from Figure 10 by referring to the tag column pattern table 90 (Figure 7). If there is no delimiter tag S, there will be many combinations of patterns that match the named entity tag column, which may lead to incorrect structural analysis. For example, as shown in Figure 11(a), if the pattern matching recognizes "Pattern 1 → Pattern 1 → Pattern 1 → Pattern 1 → Pattern 3", the correct structured data 40 is obtained, similar to Figure 8 (an example of structural analysis results in this embodiment). However, as shown in Figure 11(b), if the pattern matching recognizes "Pattern 2 → Pattern 2 → Pattern 2 → Pattern 2", incorrect structured data 40 that does not fit the context of the claim is obtained.

[0073] In the absence of the delimiter tag S, the number of possible patterns to match with the named entity tag sequence increases, making pattern matching complex and potentially leading to incorrect structural analysis. However, by introducing the delimiter tag S as in this embodiment, it becomes necessary to consider only the pattern matching of the named entity tag sequence separated by the delimiter tag S (see Figure 8). This reduces the complexity of pattern matching, enabling easy and accurate structural analysis of the document.

[0074] (3. Hardware Configuration) Next, with reference to Figure 12, the hardware configuration of the computer 30 to which the document analysis system 1 is applied will be described. As shown in Figure 12, the computer 30 is configured by connecting a control unit 31, a storage device 32, an input device 33, a display device 34, a media input / output device 35, a communication I / F unit 36, a peripheral device I / F unit 37, etc., via a bus 39. However, it is not limited to this, and various configurations can be taken as appropriate. In addition, the computer 30 may be composed of one or more computers.

[0075] The control unit 31 is composed of a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc. The control unit 31 calls programs stored in the storage device 32, ROM, recording medium (media), etc., into the work memory area on RAM and executes them, and drives and controls each part connected via the bus 39.

[0076] ROM permanently stores the boot program, BIOS, and other programs and data of the computer 30. RAM temporarily stores loaded programs and data, and also includes a work area used by the control unit 31 for various processing described later.

[0077] Furthermore, the control unit 31 executes the learning process shown in Figures 13 and 14, the analysis process shown in Figure 15, the search process shown in Figure 17, etc., according to the processing program stored in the storage device 32. The program that executes each process may be stored in advance in the storage device 32 or ROM of the computer 30, or it may be downloaded via a network or the like and stored in the storage device 32 or the like.

[0078] The storage device 32 is an HDD (hard disk drive) or the like, and stores the program executed by the control unit 31, the data necessary for program execution (pre-trained model 6, re-pre-trained model 7, named entity recognition model 8, document extraction DB 100, etc.), the OS (operating system), etc. These program codes are read by the control unit 31 as needed, transferred to RAM, and then read by the CPU for execution.

[0079] The input device 33 is, for example, a pointing device such as a keyboard, mouse, touch panel, or tablet, or a numeric keypad, and outputs the input data to the control unit 31.

[0080] The display device 34 consists of a display device such as an LCD panel or CRT monitor, and a logic circuit (such as a video adapter) for performing display processing in cooperation with the display device. The control unit 31 controls the display device to display the input display information. Alternatively, the input device 33 and the display device 34 may be integrated into a touch panel type input / output unit.

[0081] The media input / output device 35 is an input / output device for various recording media (media), such as a CD / DVD drive, and performs data input and output. The communication I / F unit 36 ​​has a communication control device, a communication port, etc., and is an interface that mediates communication with external devices connected via a network, and performs communication control.

[0082] The peripheral device interface (I / F) section 37 is a port for connecting peripheral devices to the computer 30, and the computer 30 sends and receives data with peripheral devices via the peripheral device interface (I / F) section 37. The peripheral device interface (I / F) section 37 is composed of USB, IEEE1394, etc. The connection method to peripheral devices can be wired or wireless. The bus 39 is a path that mediates the exchange of control signals, data signals, etc. between each device.

[0083] (4. Processing by Document Analysis System 1) Next, the processing of the document analysis system 1 will be described. In the document analysis system 1, the control unit 31 of the computer 30 performs a learning process (Figures 13 and 14) to learn the pre-trained model 6, the re-pre-trained model 7, and the named entity recognition model 8. The control unit 31 also uses the named entity recognition model 8 learned through the learning process to extract named entity information from the target document 2 to be analyzed, and performs the analysis process of the target document 2 (Figure 15) based on the extracted data 10.

[0084] (4-1. Learning Process) First, the learning process will be explained with reference to Figure 13. The control unit 31 (pre-training unit 14) of the computer 30 performs unsupervised pre-training using the large document set 3 and generates a pre-trained model 6 (step S11, pre-training process). The generated pre-trained model 6 is stored in the storage device 32.

[0085] Next, the control unit 31 (re-pre-training unit 15) of the computer 30 re-pre-trains the pre-trained model 6 generated in step S11 using the limited document group 4A, which is a set of documents in a specific technical field, and generates a re-pre-trained model 7 (step S12, re-pre-training process). The generated re-pre-trained model 7 is stored in the storage device 32.

[0086] Next, we will describe the actual training process for generating training data (annotation data) for training the named entity recognition model 8, and then fine-tuning the pre-trained model 6 or the re-pre-trained model 7 using the training data (annotation data) to generate the named entity recognition model 8.

[0087] First, the control unit 31 (annotation unit 16) of the computer 30 receives annotation settings for various named entity tags from the user for the limited document group 4B, which is a group of documents in a specific technical field, and generates annotated training data (annotation data 5) (step S13).

[0088] In particular, in this embodiment, the control unit 31 (annotation unit 16) of the computer 30 accepts the setting of a delimiter tag S to be attached as a named entity tag to character information that separates a sequence of named entity tags that form a certain semantic relationship (see Figure 2).

[0089] Furthermore, the control unit 31 (annotation unit 16) may accept annotation settings from the user that associate a primary named entity tag T1 with a subordinate named entity tag T2 that is dependent on the primary named entity tag T1 (see Figure 3). Also, the control unit 31 (annotation unit 16) may accept settings from the user for associated named entity tags Tr for cue words that associate named entity tags (see Figure 3).

[0090] Next, the control unit 31 (named entity recognition model learning unit 17) of the computer 30 generates a named entity recognition model 8 by fine-tuning the pre-trained model 6 generated in step S11 or the re-pre-trained model 7 generated in step S12 with the annotation data 5 (step S14, fine-tuning process). The above process generates a named entity recognition model 8. The generated named entity recognition model 8 is stored in the storage device 32.

[0091] Figure 14 is a flowchart illustrating a retraining process to further improve the learning performance of the named entity recognition model 8. The control unit 31 (document input unit 11) of the computer 30 receives input of a limited document group 4C, which is a group of documents in a specific technical field (step S31). Next, the control unit 31 (named entity recognition unit 12) of the computer 30 extracts named entity information from each document using the named entity recognition model 8 and outputs extracted data 9. In this embodiment, the extracted named entity information includes information for each named entity tag, including the delimiter tag S.

[0092] Next, the control unit 31 (re-annotation unit 18) of the computer 30 displays the extracted data 9 on the display device 34, accepts tag correction operations (correction operations for each named entity tag including the delimiter tag S) for each extracted named entity tag, and regenerates the limited document group 4C (re-annotation data 20) with the corrected named entity tags set as training data (step S33).

[0093] Then, the control unit 31 (named entity recognition model refining tuning unit 19) of the computer 30 refining the named entity recognition model 8 using the re-annotation data 20 (step S34). Steps S31 to S34 can be repeated until the desired learning performance is obtained.

[0094] (4-2. Analysis Processing) Next, the analysis process will be explained with reference to Figure 15. First, the control unit 31 (document input unit 11) of the computer 30 receives input of one or more target documents 2, which are the documents to be analyzed (step S51).

[0095] Next, the control unit 31 (named entity recognition unit 12) of the computer 30 extracts named entity information from the target document 2 using the named entity recognition model 8 and outputs extracted data 10 (step S52). In this embodiment, the extracted named entity information includes information for each named entity tag, including the delimiter tag S.

[0096] Next, the control unit 31 (structural analysis unit 13) of the computer 30 performs a structural analysis of the target document 2 based on the extracted data 10 and generates structured data 40 (step S53). Specifically, the control unit 31 (structural analysis unit 13) obtains each named entity tag sequence separated by the delimiter tag S from the delimiter tag S contained in the extracted data 10 as a named entity tag sequence (meaningful unit) that forms a single semantic relationship. Then, the control unit 31 (structural analysis unit 13) refers to the tag sequence pattern table 90 (Figure 7) to identify the pattern of each obtained named entity tag sequence and converts the information of each named entity tag sequence into structured data 40 in the data format corresponding to each pattern (see Figure 8).

[0097] Then, the control unit 31 (structural analysis unit 13) of the computer 30 stores the generated structured data 40 in the document extraction DB 100, linking it with a document ID and document attributes that uniquely identify the document (step S54).

[0098] As described above, in this embodiment, annotation settings for named entity tags are accepted for a limited document group 4B, which is a group of documents in a specific technical field, and a named entity recognition model 8 is generated using the annotated training data. In particular, in this embodiment, when setting annotations, it is possible to set a delimiter tag S, which is attached to character information that separates a sequence of named entity tags that form a certain semantic relationship, as a named entity tag. The named entity recognition model 8 extracts information of the delimiter tag S along with information of other named entity tags from the target document 2 to be analyzed. As a result, a sequence of named entity tags (a semantic unit) that forms a single meaning can be identified from the target document 2 based on the delimiter tag S, making it easier to understand and analyze the structure of technical documents.

[0099] (5. Document Search) By performing the analysis process shown in Figure 15 on numerous target documents 2, a large amount of structured data 40 is accumulated in the document extraction DB 100 (see Figure 9) for each document. Here, we will provide supplementary information on document retrieval as an example of how the document extraction DB 100 can be used.

[0100] Figure 16 is a functional block diagram showing the document search function of the document analysis system 1. The document search function mainly consists of a search condition setting unit 21, a search unit 22, and a search result display unit 23.

[0101] The search condition setting unit 21 receives input from the user, such as search keywords, numerical conditions, and unit conditions, and sets the search conditions. The search unit 22 searches the document extraction DB 100 for documents containing structured data that satisfy the search conditions. The search result display unit 23 displays information about the searched documents.

[0102] Referring to Figure 17, the search process flow will be explained. First, the control unit 31 (search condition setting unit 21) of the computer 30 receives input such as search keywords, numerical conditions, and unit conditions from the user and sets the search conditions (step S71). Next, the control unit 31 (search unit 22) searches the document extraction DB 100 for documents containing structured data that satisfy the search conditions set in step S71 (step S72). Then, the control unit 31 (search result display unit 23) displays the document search results on the display device 34 (step S73).

[0103] Preferred embodiments of the document analysis system 1, etc., according to the present invention have been described above with reference to the attached drawings, but the present invention is not limited to such examples. It will be clear to those skilled in the art that various modifications or alterations can be conceived within the scope of the technical idea disclosed herein, and these will naturally also fall within the technical scope of the present invention. [Explanation of Symbols]

[0104] 1………………Document analysis system 2………………Target Documents 3………………Large-scale document collection 4A~4C... Limited document group 5………………Annotation data 6………………Pre-trained model 7………………Re-pre-trained model 8………………Named entity recognition model 9, 10... Extracted data 11……………Document Input Section 12……………… Named entity extraction unit 13……………Structural Analysis Department 14……………Pre-learning section 15……………Pre-learning section 16……………Annotation Section 17……………Named entity recognition model learning unit 18……………Re-annotation section 19……………Named entity recognition model refining and tuning section 20……………Re-annotated data 21……………Search Criteria Setting Section 22……………Search section 23……………Search Results Display Section 30……………Computer 40……………Structured data 90……………Tag column pattern table 100…………Document Extraction Database S………………Separator tag T1……………Main Named Expression Tag T2……………Derived Unique Entity Tag Tr... Related named entity tags

Claims

1. An annotation unit that accepts annotation settings for named entity tags for a limited set of documents which is a set of documents in a specific technical field, and generates annotated training data, wherein the annotation unit is capable of setting delimiter tags to be attached to character information that separates a sequence of named entity tags which form a certain semantic relationship as the named entity tags, A named entity recognition model learning unit generates a named entity recognition model using the training data which includes the delimiter tag as a named entity tag. A document input unit that inputs documents in the aforementioned technical field as the target for analysis, A named entity recognition unit that uses the named entity recognition model to extract information of delimiter tags along with information of other named entity tags from the document to be analyzed, A document analysis system characterized by comprising the following features.

2. The document further comprises a structural analysis unit that, by referring to the delimiter tags extracted by the named entity recognition unit, obtains each sequence of named entity tags separated by the delimiter tags as a sequence of named entity tags forming a single semantic relationship, and performs structural analysis of the document. The document analysis system according to claim 1, characterized in that it is the same as described in claim 1.

3. The annotation unit accepts annotation settings for named entity tags having a master-slave relationship, and includes the named entity tags having a master-slave relationship in the training data to be used for learning and extraction of named entity recognition. The document analysis system according to claim 1, characterized in that it is the same as described in claim 1.

4. The annotation unit accepts annotation settings for related named entity tags for cue words that associate named entity tags, and includes the related named entity tags in the training data to be used for named entity recognition learning and extraction. The document analysis system according to claim 1, characterized in that it is the same as described in claim 1.

5. The structural analysis unit refers to a tag column pattern table that stores pre-associated patterns of named entity tag columns with data formats of structured data, identifies the pattern of each named entity tag column by pattern matching, and converts the information of each named entity tag column into structured data in the corresponding data format. The document analysis system according to claim 2, characterized in that it is the same as described in claim 2.

6. The tag column pattern table holds multiple patterns with different word orders for named entity tag columns that represent the same meaning. The document analysis system according to claim 5, characterized in that it is a document analysis system.

7. The named entity tag includes an element tag, a numeric tag, and a unit tag. The document analysis system according to claim 1, characterized in that it is the same as described in claim 1.

8. The aforementioned limited set of documents and the documents to be analyzed are intellectual property-related documents and documents. The document analysis system according to claim 1, characterized in that it is the same as described in claim 1.

9. Computers An annotation process for a limited set of documents, which is a set of documents in a specific technical field, which accepts the setting of annotation settings for named entity tags and generates annotated training data, wherein the annotation process allows setting of delimiter tags to be attached to character information that separates a sequence of named entity tags that form a certain semantic relationship as the named entity tag, A named entity recognition model learning process that generates a named entity recognition model using the training data which includes the delimiter tag as a named entity tag, A document input process in which documents in the aforementioned technical field are input as the target of analysis, A named entity recognition step is performed using the named entity recognition model to extract information of delimiter tags along with information of other named entity tags from the document to be analyzed. A document analysis method characterized by performing the following.

10. Computers, An annotation unit that accepts annotation settings for named entity tags for a limited set of documents which is a set of documents in a specific technical field, and generates annotated training data, wherein the annotation unit is capable of setting delimiter tags to be attached to character information that separates a sequence of named entity tags which form a certain semantic relationship as the named entity tags, A named entity recognition model learning unit generates a named entity recognition model using the training data which includes the delimiter tag as a named entity tag. A document input unit that inputs documents in the aforementioned technical field as the target for analysis, A named entity recognition unit that uses the named entity recognition model to extract information of delimiter tags along with information of other named entity tags from the document to be analyzed, A program characterized by its ability to function in a certain way.

Citation Information

Patent Citations

  • Text punctuation detection method, computer equipment and storage medium

    CN114298032A

  • Information processing apparatus and program

    JP2010092169A

  • Information processing method, natural language processing method, and information processing apparatus

    JP2020098594A

  • JPP7682862B

  • Speech recognition text processing method and apparatus, device, storage medium, and program product

    US20230289514A1