Document analysis device and document analysis program

The document analysis device automates the classification and comparison of software development requirements to efficiently utilize past assets, addressing manual inefficiencies and granularity issues in software development analysis.

JP7839708B2Active Publication Date: 2026-04-02ASTEMO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for analyzing software development requirements struggle with efficiently identifying and utilizing past software assets due to manual processes, complexity in handling diverse customer requirements, and difficulty in accurately extracting differences between new and past software asset descriptions, especially when customers have different description granularities.

Method used

A document analysis device and program that classify requirements into groups, extract topics from these groups, compare and identify differences between documents, and output analysis results to facilitate efficient use of past software assets in new development.

Benefits of technology

Enhances the efficiency of software development by accurately identifying and utilizing past software assets, improving the analysis process through automated classification and comparison of document requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839708000001
    Figure 0007839708000001
  • Figure 0007839708000002
    Figure 0007839708000002
  • Figure 0007839708000003
    Figure 0007839708000003
Patent Text Reader

Abstract

To enable past software assets to be efficiently used in software development and increase efficiency in the software development.SOLUTION: A document analysis device comprises: a group classification unit which discriminates requirements included in a first document to be analyzed and classifies the discriminated requirements into a plurality of groups; a topic extraction unit which extracts terms related to the requirements classified into the plurality of groups as a topic; a topic difference extraction unit which compares a topic included in a group of an analyzed second document different from the first document with a topic included in the group of the first document and extracts the difference; and an analysis result output unit which outputs the analysis results indicative of the analysis results including the difference to the outside.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a document analysis apparatus and a document analysis program.

Background Art

[0002] In software development, analysis of a document that describes requirements specifications, etc. of software to be newly developed presented by a new customer is performed. According to the analysis result, it is generally known to examine whether it is possible to reuse past software assets, and if it is possible to reuse, to use it.

[0003] At present, such analysis and examination are mainly performed manually by developers, etc. based on the knowledge and experience of developers, etc. However, when customer requirements are diverse, the analysis work becomes complicated, and when the number of past software assets increases, it also takes a very long time to search for them. As a result, it becomes difficult to effectively use past software assets.

[0004] A system that supports such analysis and examination by a computer is also known, for example, from Patent Document 1. Patent Document 1 discloses a technique for comparing paragraphs and sections of a document by a computer to determine the similarity between two documents. With this technique, it is only possible to determine whether a new document contains a new paragraph in comparison with a past document, and it is difficult to use it for determining whether past software assets can be used in new software development.

[0005] In addition, when a new customer is different from a customer of past software assets, the description granularity (process, method, standard, etc.) of requirements may be different, and it is difficult to accurately extract the difference between the requirements of the new customer and the requirements in past software assets. As a result, there is a problem that it becomes difficult to identify past software assets that can be used.

Prior Art Documents

Patent Documents

[0006] [Patent Document 1] Japanese Patent Publication No. 2015-219799 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] This disclosure was made in view of the above-mentioned issues and provides a document analysis device and a document analysis program that enable the efficient use of past software assets in software development and improve the efficiency of software development. [Means for solving the problem]

[0008] To solve the above problems, the document analysis device according to this disclosure is characterized by comprising: a group classification unit that identifies requirements contained in a first document to be analyzed and classifies them into a plurality of groups; a topic extraction unit that extracts terms related to the requirements classified into the plurality of groups as topics; a topic difference extraction unit that compares topics contained in a group of a second document that has been analyzed but is different from the first document with topics contained in the group of the first document and extracts the difference; and an analysis result output unit that outputs an analysis result showing the results of the analysis including the difference to an external source. [Effects of the Invention]

[0009] The document analysis device described herein enables the efficient use of past software assets in software development, thereby improving the efficiency of software development. It also provides a document analysis device and a document analysis program. [Brief explanation of the drawing]

[0010] [Figure 1A] This is a schematic diagram illustrating the document analysis device 200 and user terminal 100 according to the first embodiment. [Figure 1B] This is a block diagram illustrating in more detail the configuration of the document analysis device 200 according to the first embodiment. [Figure 2] This is a schematic diagram illustrating the analysis process of a new request document in the document analysis device 200 according to the first embodiment. [Figure 3] This is a schematic diagram illustrating a document analysis device 200 and a user terminal 100 according to a second embodiment. [Figure 4A] This flowchart illustrates the procedures for the new request document analysis process, the display control process for the analysis results of the new request document, the group classification verification process, the topic extraction verification process, and the score calculation process in the document analysis device 200 of the second embodiment. [Figure 4B] This flowchart illustrates the procedures for the new request document analysis process, the display control process for the analysis results of the new request document, the group classification verification process, the topic extraction verification process, and the score calculation process in the document analysis device 200 of the second embodiment. [Figure 5] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 6] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 7] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 8] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 9] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 10] An example of a screen display showing the comparison result between a new request document and a past request document in the user terminal 100 of the second embodiment will be described. [Figure 11] This flowchart illustrates an example of the procedure for controlling the display of analysis results in the user terminal 100 of the second embodiment. [Figure 12A] This is a flowchart for explaining an example of the procedure of display control processing of analysis results in the user terminal 100 according to the second embodiment. [Figure 12B] This is a flowchart for explaining an example of the procedure of display control processing of analysis results in the user terminal 100 according to the second embodiment. [Figure 13] This is a flowchart for explaining an example of the procedure for update control of display of analysis results according to the second embodiment. [Figure 14] This is a flowchart for explaining an example of the procedure for update control of display of analysis results according to the second embodiment. [Figure 15] This is a flowchart for explaining the procedure of update control for updating the analysis results of a new claim document according to the second embodiment.

Embodiments for Carrying Out the Invention

[0011] Hereinafter, this embodiment will be described with reference to the accompanying drawings. In the accompanying drawings, functionally identical elements may sometimes be denoted by the same number. Note that the accompanying drawings show embodiments and implementation examples in accordance with the principles of the present disclosure, but these are for the purpose of understanding the present disclosure and are not used to limit the interpretation of the present disclosure in any way. The description in this specification is merely a typical example and does not limit the scope or application examples of the claims of the present disclosure in any sense.

[0012] In this embodiment, although the description is made in sufficient detail for those skilled in the art to implement the present disclosure, other implementations and forms are also possible, and it is necessary to understand that changes in configuration and structure and replacement of various elements can be made without departing from the scope and spirit of the technical idea of the present disclosure. Therefore, the following description should not be construed as being limited thereto.

[0013] [First Embodiment] Referring to Figure 1A, the document analysis device 200 and user terminal 100 according to the first embodiment will be described. The document analysis device 200 of the first embodiment is connected to the user terminal 100 and receives documents related to the design specifications of newly developed software (hereinafter referred to as "new request documents" or "first documents") from the user terminal 100.

[0014] The document analysis device 200 analyzes the new request document and, based on the analysis results, identifies documents that have similarities with the new request document from among past request documents that have already been analyzed and whose analysis results have been stored (hereinafter referred to as "past request documents" or "second documents"). The document analysis device 200 then identifies the similarities, differences, and new features between the identified related past request documents and the new request document and presents them to the user terminal 100. The user (software developer) of the user terminal 100 can look at the presented past request documents and the information on their similarities, differences, and new features and determine whether the past software assets related to those past request documents can be used for the development of new software related to the new request document.

[0015] The user terminal 100 can be configured as a general-purpose personal computer or the like, and may include, for example, a CPU 101, ROM 102, RAM 103, hard disk drive 104, input / output control unit 105, communication control unit 106, display control unit 107, input device 108, and display 109. The storage device such as the hard disk drive 104 stores a user interface application that constitutes part of the document analysis program for the operation of the document analysis device 200 in this embodiment. Input from the user, such as various instructions and editing operations, is received from the input device 108. The display 109 may display the execution screen of the user interface application.

[0016] The document analysis device 200 can similarly be configured using a general-purpose personal computer, for example, comprising a CPU 201, ROM 202, RAM 203, hard disk drive 204, input / output control unit 205, communication control unit 206, and display control unit 207. The storage device, such as the hard disk drive 204, stores a document analysis program for the operation of the document analysis device 200 in this embodiment. Although not shown in Figure 1A, the document analysis device 200 may also include an input device operated by the administrator of the document analysis device 200 and a display for confirming the analysis operation.

[0017] The document analysis program is implemented in the document analysis device 200 as a document analysis processing unit 211, a document analysis model generation unit 212, a document analysis result management unit 213, and a document analysis result input / output unit 214. The document analysis processing unit 211 receives data of a new request document and performs various analyses related to the new request document. The document analysis model generation unit 212 generates document analysis models (request classification model, named entity recognition model) used for analysis in the document analysis processing unit 211.

[0018] The document analysis result management unit 213 is responsible for managing data related to the analysis results of newly requested documents, data related to the analysis results of past requested documents, and various other data used in the analysis. The document analysis result input / output unit 214 generates display data for displaying the analysis results of newly requested documents on the user terminal 100 and outputs it to the user terminal 100, and also has the function of modifying this display data by receiving various inputs from the user terminal 100, etc.

[0019] As shown in Figure 1B, the document analysis processing unit 211 further includes, as an example, a group classification unit 2111, a topic extraction unit 2112, a topic difference extraction unit 2113, and a new request document creation unit 2114. The group classification unit 2111 has the function of identifying the requirements contained in the new request document to be analyzed and classifying them into multiple groups. The topic extraction unit 2112 has the role of extracting terms related to the terms (keywords) contained in the requirements classified into multiple groups as topics. The topic difference extraction unit 2113 has the role of comparing the topics contained in one group of previously analyzed request documents with the topics contained in the group of the new request document and extracting the differences. The new request document creation unit 2114 has the function of generating a new request document that includes the results of the difference extraction. The topic difference extraction unit 2113 may also have the function of calculating the topic match rate and vector similarity calculated based on the differences.

[0020] The document analysis model generation unit 212 generates a request classification model 2121 used for classification processing in the group classification unit 2111 of the document analysis processing unit 211, and also generates a named entity recognition model 2122 used for topic extraction in the topic extraction unit 2112. The request classification model 2121 and the named entity recognition model 2122 together constitute the document analysis model. The document analysis model can be updated as appropriate using natural language processing and machine learning techniques. The topic extraction unit 2112 can be composed of either the multi-label request classification model 2121' or the named entity recognition model 2122, or both. The multi-label request classification model 2121' is a model that gives the topic extraction unit 2112 the ability to extract multiple topics. On the other hand, the request classification model 2121 is limited to a single label (group). The request classification models 2121, 2121', and the named entity recognition model 2122 can be implemented as different models (software).

[0021] Note that the named entity recognition model 2122 may be omitted depending on the circumstances. Also, the requirements classification model 2121 and the named entity recognition model 2122 may be generated as separate models depending on the group. For example, if there are 10 groups, 10 named entity recognition models 2122 and 10 requirements classification models 2121 may be generated.

[0022] The document analysis result management unit 213 further includes, as an example, a new request document management unit 2131, a past request document management unit 2132, a topic data management unit 2133, a group data management unit 2134, a document analysis result data management unit 2135, and a document analysis result update control unit 2136.

[0023] The new request document management unit 2131 is responsible for managing new request documents, specifically managing, for example, the original data of new request documents, the classification results of the group classification unit 2111 for new request documents, the extraction results of the topic extraction unit 2112, and other data related to new request documents. The past request document management unit 2132 is responsible for managing past request documents, specifically managing the original data of past request documents, the classification results of the group classification unit 2111 for past request documents, the extraction results of the topic extraction unit 2112, and other data related to past request documents.

[0024] The topic data management unit 2133 is used in the topic extraction process in the topic extraction unit 2112 and manages data related to topics using a database. The group data management unit 2134 is used in the classification process in the group classification unit 2111 and manages data related to groups using a database. The document analysis result data management unit 2135 is responsible for managing analysis result data as a result of analyzing new requested documents. The document analysis result update control unit 2136 is responsible for update control to update the analysis result data.

[0025] Referring to Figure 2, the analysis process of a new request document in the document analysis device 200 will be explained. As shown in the upper left of Figure 2, a new request document contains multiple requirements, New Req-i. Similarly, past request documents also contain multiple requirements, Old Req-i. Here, "requirements" are sentences that express various requirements for the development of a system or service in a single document. Requirements may be a single sentence (a sentence with only one period) or multiple sentences.

[0026] The requirements (New Req-i) in new request documents are classified into multiple groups in the group classification unit 2111 according to the requirement classification model and group database, depending on their content. Examples of these groups include "object detection," "diagnosis," and "sensor performance," as shown in Figure 2. The requirements (Old Req-i) in past request documents are similarly classified into multiple groups.

[0027] New Req-i, which is classified into one of several groups, is subjected to topic extraction processing in the topic extraction unit 2112, and terms contained in New Req-i are extracted as topics. The results of the group classification and topic extraction are stored in the new request document management unit 2131.

[0028] Furthermore, the extracted topic expressions (terms) are converted to other terms as appropriate according to the topic database (for example, "driving lane" is changed to "white line"). In other words, a "topic" may include not only the term itself contained in the original text of a new or past request document, but also related terms (e.g., higher-level terms, lower-level terms, synonyms, etc.). Past request documents are also subject to topic extraction, and the results of the extraction are stored in the past request document management unit 2132.

[0029] When the group classification and topic extraction results of a new request document are stored in the new request document management unit 2131, the topic difference extraction unit 2113 of the document analysis processing unit 211 performs a comparison of topics between the past request documents stored in the past request document management unit 2132 and the corresponding groups, and extracts the topic differences between the two (topics that match between the new request document and the past request documents, topics that are missing in the new request document, and new topics in the new request document). This extraction is performed between the new request document and multiple past request documents. The user of the user terminal 100 can look at the results of this extraction to identify the past request document that is closest to the new request document, and can use the past software assets related to that past request document for software development related to the new request document.

[0030] The topic difference extraction unit 2113 may extract topic differences between groups having the same or related group names, but is not limited to this; it may also be capable of extracting topic differences between groups having different group names. Furthermore, the comparative analysis performed by the topic difference extraction unit 2113 is not limited to two groups; the comparative analysis can be performed on any group as long as the topics are comparable. For example, the requirement "New Req" in a new request document may be compared with the group in the past request document being compared.

[0031] As described above, according to the document analysis device 200 of the first embodiment, the requirements contained in the document are classified into groups, and within each group, the terms in the requirements are extracted as topics. Then, by comparing the topics for each group, the degree of similarity with past request documents is determined. This makes it possible to accurately identify past request documents that are similar to the new request document.

[0032] [Second Embodiment] Next, with reference to Figure 3, the document analysis device 200 of the second embodiment will be described. Similar to the first embodiment, the document analysis device 200 of the second embodiment is connected to a user terminal 100 and receives documents related to the design specifications of newly developed software (hereinafter referred to as "new request documents" or "first documents") from the user terminal 100. However, this document analysis device of the second embodiment differs from the first embodiment in that it is equipped with a document analysis reliability calculation unit 215 that calculates the reliability of the document analysis results, and the document analysis model generation unit 212 is equipped with a vector similarity calculation model generation unit 2123. By calculating the reliability of the document analysis results and presenting it to the user terminal 100, it becomes possible to make more accurate judgments about the document analysis results.

[0033] The document analysis reliability calculation unit 215 includes, as an example, a topic matching rate calculation unit 2151, a vector similarity calculation unit 2152, and a topic matching rate / vector similarity difference calculation unit 2153. The topic matching rate calculation unit 2151 has the function of calculating the topic matching rate, which indicates the degree of topic matching within a group between a new request document and a past request document. The vector similarity calculation unit 2152 has the function of calculating the similarity of topics within a group between a new request document and a past request document as a vector similarity such as cosine similarity. The topic matching rate / vector similarity difference calculation unit 2153 has the function of calculating the difference between the topic matching rate calculated by the topic matching rate calculation unit 2151 and the vector similarity calculated by the vector similarity calculation unit 2152, and comparing this difference with a threshold. The reliability of the document analysis can be determined according to the difference between this difference and the threshold.

[0034] Next, referring to the flowcharts in Figures 4A and 4B, the procedures for the new request document analysis process, the display control process for the analysis results of the new request document, the group classification verification process, the topic extraction verification process, and the score calculation process in the document analysis device 200 of the second embodiment will be described.

[0035] In the new request document analysis process, first, group classification of the requirements contained in the new request document is performed (step S11). Then, terms contained in the classified requirements are extracted as topics (steps S12, S13). In step S12, topic extraction from the new request document is performed according to a named entity recognition model, and in step S13, the terms related to the topics extracted from the new request document are converted to other terms according to the topic database. Based on the results of the topic extraction in steps S12 and S13, a new request document with group classification and topic extraction is created (step S14).

[0036] Next, group information of past request documents is obtained from the past request document management unit 2132, along with topic extraction information of past request documents (steps S15, S16). Then, past request documents are created with topic replacements on a group basis as needed (step S17). The newly generated request documents and past request documents are then subjected to topic difference extraction on a group basis (step S18).

[0037] When the topic differences between groups of new request documents and past request documents are extracted, the topic matching rate between the groups is calculated based on these differences (step S21). Furthermore, the average value of vector similarity is calculated for each group in the new request documents (step S22), and information regarding the average value of vector similarity for each group in the past request documents is read and obtained from the past request document management unit 2132 (step S23). Then, the difference in vector similarity between groups of new request documents and past request documents is calculated (step S24). Furthermore, the difference between the topic matching rate and vector similarity between the new request documents and past request documents is calculated, and the reliability of the document analysis is determined based on this (step S25). Then, an analysis is performed according to the results of the above calculations, and the analysis results are displayed on the user terminal 100 (step S26).

[0038] Referring to Figures 5 and 6, an example of the screen display for the comparison results between a new request document and a past request document on the user terminal 100 will be explained. Figure 5 is an overview of the screen, and Figure 6 shows a detailed example. This screen includes, as an example, the analysis / comparison target designation display screen 2, the analysis result list display / analysis result detail selection screen 3, and the analysis result detail display / edit screen 4. The analysis / comparison target designation display screen 2 includes a screen for designating (selecting) a new request document as the analysis / comparison target, a screen for designating (selecting) a past request document to be compared with the new request document, and a screen for selecting the analysis scores for both.

[0039] The Analysis Results List Display / Analysis Results Details Selection Screen 3 displays a list of analysis results for a new request document and allows users to selectively view the details of those analysis results. The Analysis Results List Display / Analysis Results Details Selection Screen 3 further includes, as an example, a Classification Confidence Score Table 10 and a Topic Extraction Confidence Score Table 11. The Classification Confidence Score Table 10 displays the confidence level of the group classification decisions as a score. The Topic Extraction Confidence Score Table 11 displays the confidence level of the topic extraction process in the Topic Extraction Unit 2112 as a score.

[0040] The detailed analysis results display / edit screen 4 includes, as an example, a new request document display / edit screen 12, a past request document display / edit screen 13, and a topic difference display screen 14. The new request document display / edit screen 12 is a screen for displaying and editing the analysis results for a new request document. The past request document display / edit screen 13 is a screen for displaying and editing the analysis results for a past request document that is compared with the new request document. The topic difference display screen 14 is a screen that displays the differences between the new request screen and the past request screen, as well as various factors related to those differences.

[0041] As shown in Figure 6, the new request document display / edit screen 12 includes a group name display field 12A as a result of group classification of the new request document, a source text display field 12B that displays the source data of the new request document, and a topic / source word display field 12C that shows the correspondence between extracted topics and corresponding words in the source text. Below fields 12A to 12C, icons may be displayed to indicate editing, saving, and completion of analysis for this data. Figure 7 shows a specific example of the display in fields 12A to 12C. In the source text display field 12B, it is possible to indicate the location of a topic in the source text using symbols (<>, etc.). 7 In the topic / source word display field 12C, the relationship between the topic and the corresponding location in the source text can be understood, and the user can also edit the topic expression on the user terminal 100. Furthermore, it is possible to check the topic string and the corresponding location in the source text and register the term in a topic database, etc. Note that fields 12B and 12C may be combined and displayed in a single field as shown in Figure 8.

[0042] The past request document display / editing screen 13 includes a group name display field 13A as a result of group classification of past request documents to be compared with the new request document, a source text display field 13B that displays the source data of the past request document, and a topic / source word display field 13C that shows the correspondence between extracted topics and corresponding words in the source text. Below fields 13A to 13C, icons for instructing editing and saving of this data may be displayed. Figure 7 shows a specific example of the display in fields 13A to 13C.

[0043] The detailed analysis results display / editing screen 4 includes a re-analysis start instruction button 15A, a Prev button 15B, and a Next button 15C. The re-analysis start instruction button 15A instructs the system to re-run the analysis on the new request documents and past request documents currently displayed in columns 12 and 13. buttonThe Prev button 15B and Next button 15C are buttons used to switch the display of the filtered list of analysis results on the Analysis / Comparison Target Selection Display Screen 2. The Prev button 15B and Next button 15C are buttons used to switch the past request documents displayed on the Past Request Document Display / Edit Screen 13. When these are pressed, the new request documents, past request documents, and others displayed on the Analysis / Comparison Target Selection Display Screen 2 are switched, and the new analysis results are displayed on the Topic Difference Display Screen 14.

[0044] The topic difference display screen 14 is a screen for displaying the differences in topics between the new request document displayed on screen 12 and the past request document displayed on screen 13, in groups. Specifically, topics common to both documents are displayed as "common topics," topics that exist only in the past request document and are missing (absent) in the new request document are displayed as "missing topics," and topics that appear only in the new request document are displayed as "new topics." Figure 7 shows a specific example of the display on the topic difference display screen 14.

[0045] Referring to Figure 9, a first modified example of the screen display showing the comparison results between a new request document and a past request document on the user terminal 100 will be explained. The screen display in Figure 9 differs from the display example in Figure 6 in that the new request document display / editing screen 12 and the past request document display / editing screen 13 are equipped with topic string / original text position display fields 12D and 13D that show the extracted topic string and the position in the original text where the term corresponding to that topic appears. By showing the topic string and the position in the original text where the corresponding term appears, it becomes easier to compare new request documents and past request documents.

[0046] Referring to Figure 10, a second modified example of the screen display showing the comparison results between a new request document and a past request document on the user terminal 100 will be described. The screen display in Figure 10 differs from the display example in Figure 6 in that multiple sets of new request document display / editing screens 12 and past request document display / editing screens 13 are displayed in parallel. As a result, the comparison results of multiple past request documents are displayed on a single screen. The user of the user terminal 100 can more easily determine which of the multiple past request documents has a high degree of similarity to the new request document.

[0047] Referring to the flowcharts in Figures 11 to 12B, an example of the procedure for controlling the display of analysis results on the user terminal 100 will be explained. First, steps S31 and S32 are executed, which are procedures for sorting the list of analysis results obtained by comparing and analyzing a new request document with multiple past request documents. Step S31, as an example, compares the vector similarity scores among multiple past request documents and sorts the analysis results in descending order of vector similarity scores. Step S32, as an example, compares the degree of agreement in group classification among multiple past request documents and sorts the analysis results in descending order of agreement (see Figure 12B). In addition, in step S31, as shown in Figure 12A, the analysis results can be sorted in ascending order of vector similarity scores (step S31A), or the analysis results can be sorted in descending order according to the difference score between the topic agreement rate and the vector similarity score (step S31B). Furthermore, the analysis results can be sorted in ascending order by topic agreement rate (step S31C), and also sorted in descending order according to the difference score between vector similarity and topic agreement rate (step S31D).

[0048] In step S33, it is determined whether or not an instruction to end the display of analysis results has been issued. If it has been issued, (Y) the procedure in Figure 11 is completed; otherwise, (N) the process proceeds to step S34.

[0049] In step S34, data selection and filtering are performed based on the information of the new request document designated as the target of analysis. The target of analysis may be specified, for example, by specifying the document name, group name, or topic name. In the subsequent step S35, data selection and filtering are performed based on the information of the past request document designated as the comparison target. The target of analysis may be specified, for example, by specifying the document name, group name, or topic name of the past request document.

[0050] In step S36, it is determined whether or not a group has been specified in the specification of the analysis target. If a group has been specified, (N) the process proceeds to step S37; if no group has been specified, (Y) the process proceeds to step S38.

[0051] In step S37, according to the specified group, the grouping results for that group, the original text of that group, the topic extraction results within that group, and the differences in topics between the extracted topics and the corresponding group of the past request documents being compared are displayed.

[0052] On the other hand, in step S38, according to the specified new request document, the grouping results for each of the multiple groups contained in the new request document, the original text of the multiple groups, the topic extraction results for each of the multiple groups, and the differences in topics between the extracted topics and the corresponding groups of the past request document being compared are displayed. The above display control procedure continues until the instruction to end the display of analysis results is issued (step S33).

[0053] Next, with reference to Figure 13, an example of a procedure for controlling the update of the display of analysis results will be described. First, if a re-analysis start command is issued by the re-analysis start command button 15A or the like (step S51, Y), the procedures in Figures 4A and 4B are executed for the new request document and past request document currently displayed on the screen, and the procedure in Figure 13 is completed. On the other hand, if a command is issued to update the display of analysis results (step S51, N), the data of the new request document to be the target of the new analysis is displayed, for example, on the new request document display / edit screen 12 (step S52).

[0054] Next, it is determined whether or not a change to the group to be analyzed is necessary (step S53), and if necessary (Y), a group change flow is performed to change the group (step S54). Also, it is determined whether or not a change to the topic to be analyzed is necessary (step S55), and if necessary, a topic change flow is performed to change the topic to be analyzed (step S56). In this way, the update control of the analysis targets is completed, and when the re-analysis start instruction button 15A is pressed, the analysis process is executed in the same manner.

[0055] The flowchart on the left side of Figure 14 shows an example of the detailed procedure for the group change flow (step S54). When a group change is instructed, the analysis / comparison target designation display screen 2 displays a list of groups included in the new request document to be analyzed (step S54A). The user on user terminal 100 looks at this list of groups and determines whether there is a group in the list that they would like to be considered for the next analysis (step S54B). If there is a group in the list that is a candidate for the next analysis (Y), that group is selected from the group list (step S54C). If no candidate group is found (N), the new group name is entered in the search box (not shown) to search and identify the corresponding group (step S54D). Once the group to be analyzed next has been identified, the editing status flag, which indicates whether or not the corresponding new request document has been edited, is set to "TRUE". (Step S54E) .

[0056] Furthermore, the flowchart on the right side of Figure 14 shows an example of the detailed procedure for the topic change flow (step S56). When a topic change is instructed, on the analysis / comparison target designation display screen 2, the topic to be changed is deleted from the topics included in the new request document to be analyzed (step S56A), and the position of the topic within the new request document is selected (step S56B), and a list of topics corresponding to that position is displayed (step S56C). The user on user terminal 100 looks at the list and determines whether or not there is a topic in the list that is a candidate for determination (step S56D). If there is a candidate topic (Y), the candidate is selected from the list of topics (step S56E). If there is no candidate topic (N), the new topic name is entered in the search box (not shown) to search and identify the corresponding topic (step S56F). Once the topic to be analyzed next is identified, the edit flag indicating that the corresponding new request document is editable is set to "TRUE".

[0057] Next, referring to the flowchart in Figure 15, the update control procedure for updating the analysis results of a new request document will be explained. First, when the new request document management unit 2131 and the past request document management unit 2132 receive and acquire the latest new and past request documents updated by the user (step S61), it is determined whether or not there is a request to update the document analysis results for that new request document (step S62). If there is no update request, the operation ends (N), but if there is an update request (Y), it is determined whether the re-analysis requirement flag is set to "TRUE" (step S63). If it is TRUE, the document analysis model is updated (re-trained) in the document analysis model generation unit 212 (step S64), and the re-analysis of the new request document by that document analysis model is performed (steps S65-S69). Specifically, in step S66, if the flag indicating whether or not the analysis of the new request document has been confirmed is "FALSE" (analysis not yet confirmed), the procedure in Figure 4A (steps S11-S18: New Request Document Analysis Flow (1)) is executed. If the analysis of the new request document has been confirmed and the document analysis confirmation flag is set to "TRUE" (N), steps S11 to S18 are omitted, and the procedure shown in Figure 4B (steps S21 to S26: new request document analysis flow (2), (3)) is executed.

[0058] The embodiments have been described above, but it is also possible to employ the following document analysis methods. (1) The confidence coefficient R is set by the number of matching customer (=document issuer) values ​​between new request documents and past request documents. NCU This can be multiplied by a numerical value indicating inter-document similarity, topic match rate, etc., to recalculate the confidence score. This is based on the fact that the greater the number of document issuer matches between new and past requested documents, the higher the confidence of the analysis results. (2) The confidence coefficient R is set by the number of matching request groups between new request documents and past request documents. NRG This is multiplied by a numerical value indicating inter-document similarity or topic match rate to recalculate the confidence score. This is based on the principle that the more frequently the same group appears, the more reliable the analysis results become. (3) The confidence coefficient R corresponding to the ratio of the total number of requests for new documents (M) to the total number of requests for past documents (N)RNR The confidence score is recalculated by multiplying this by a numerical value indicating inter-document similarity or topic match rate. This is based on the fact that the closer the ratio of the total number of requests for new documents (M) to the total number of requests for past documents (N) is to 1, the more reliable the analysis results become.

[0059] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. The embodiments described above are described in detail for the purpose of clearly illustrating the present invention, and are not necessarily limited to those having all the configurations described. Furthermore, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace parts of the configuration of each embodiment with other configurations.

[0060] Furthermore, each of the above configurations, functions, processing units, and processing means may be implemented in hardware by designing some or all of them as integrated circuits. Furthermore, each of the above configurations and functions may be implemented in software by having the processor interpret and execute programs that implement each function. Information such as programs, tables, and files that implement each function may be stored in a memory, hard disk, or SSD recording device, or in a recording medium such as an IC card, SD card, or DVD. [Explanation of Symbols]

[0061] 2…Analysis / Comparison Target Selection Display Screen 3…Analysis Results List Display / Analysis Results Details Selection Screen 4…Detailed Analysis Results Display / Edit Screen 10…Classification Confidence Score Table 11…Topic Extraction Confidence Score Table 12… New Request Document Display / Edit Screen 13…Past Request Document Display / Edit Screen 14…Topic difference display screen 15A... Button to start reanalysis 15B... Prev button 15C...Next button 100...User terminals 104... Hard disk drive 105…Input / Output Control Unit 106...Communication Control Unit 107...Display Control Unit 108…Input devices 109…Display 200... Document analysis device 204... Hard disk drive 205… Input / Output Control Unit 206...Communication Control Unit 207...Display Control Unit 211…Document Analysis Processing Unit 212...Document Analysis Model Generation Unit 213…Document Analysis Results Management Department 214...Document Analysis Result Input / Output Unit 215...Document Analysis Confidence Calculation Unit 2111... Group Classification Department 2112... Topic extraction unit 2113...Topic difference extraction unit 2114... New Request Document Creation Department 2123... Vector Similarity Calculation Model Generation Unit 2131... New Request Document Management Department 2132…Past request document management department 2133... Topic Data Management Department 2134... Group Data Management Department 2135...Document Analysis Results Data Management Department 2136...Document Analysis Result Update Control Unit 2151...Topic Match Rate Calculation Unit 2152... Vector Similarity Calculation Unit 2153...Topic Match Rate / Vector Similarity Difference Calculation Unit

Claims

1. A group classification unit identifies the requirements contained in the first document to be analyzed, classifies them into multiple groups, and assigns group names to them. A topic extraction unit that extracts terms related to the requirements classified into the aforementioned multiple groups as topics, A topic difference extraction unit compares topics in a second document that has been analyzed and is different from the first document, which have the same group name as the group in the first document, with topics in the group in the first document, and extracts the differences. An analysis result output unit that outputs analysis results showing the results of the analysis including the aforementioned difference to an external source. Equipped with, The topic extraction unit extracts the terms themselves included in the requirements as topics, and also extracts at least one of the superordinate conceptual terms, subordinate conceptual terms, or synonyms of the terms included in the requirements as topics, considering them as terms related to the requirements. A document analysis device characterized by the following features.

2. The system further includes an analysis reliability calculation unit for calculating the reliability of the analysis of the first document, The aforementioned analysis confidence calculation unit, A topic matching rate calculation unit that calculates a topic matching rate by comparing the topics included in the requirements of the first document with the topics included in the requirements of the second document, A vector similarity calculation unit calculates the vector similarity between the topic included in the requirements of the first document and the topic included in the requirements of the second document. A topic match rate / vector similarity difference calculation unit calculates the difference between the topic match rate and the vector similarity and calculates the reliability of the analysis of the first document. Furthermore, The analysis confidence calculation unit presents the confidence level to the user terminal. The document analysis apparatus according to claim 1.

3. The document analysis apparatus according to claim 1, wherein the topic difference extraction unit identifies common topics included in both the first and second documents, missing topics that are missing in the first document, and new topics that exist only in the first document, according to the differences between the topics of the first document and the topics of the second document.

4. The document analysis apparatus according to claim 1, wherein the topic extraction unit is configured to convert the terms extracted as topics into other terms according to a database.

5. The document analysis apparatus according to claim 1, wherein the analysis result output unit outputs the results of the classification by the group classification unit and the topics extracted by the topic extraction unit, and makes the topics editable on an external device.

6. The steps include identifying the requirements contained in the first document to be analyzed, classifying them into multiple groups, and assigning group names, A step of extracting terms related to the requirements classified into the aforementioned multiple groups as topics, The steps include comparing topics in a second document that has been analyzed but is different from the first document, that are included in a group having the same group name as the group in the first document, with topics in the group in the first document, and extracting the differences; The steps include outputting the analysis results, including the aforementioned difference, to an external source. It is configured to have a computer execute it, In the step of extracting terms related to the requirements classified into the aforementioned multiple groups as topics, the terms included in the requirements themselves are extracted as topics, and at least one of the superordinate conceptual terms, subordinate conceptual terms, or synonyms of the terms included in the requirements is considered a term related to the requirements and extracted as a topic. A document analysis program.

7. The method further includes a step of calculating the reliability of the analysis of the first document, The step of calculating the confidence level is: A step of calculating a topic matching rate by comparing the topics included in the requirements of the first document with the topics included in the requirements of the second document, A step of calculating the vector similarity between the topic included in the requirements of the first document and the topic included in the requirements of the second document, The steps include calculating the difference between the topic match rate and the vector similarity, and calculating the confidence level of the analysis of the first document, The steps include presenting the aforementioned reliability level to the user terminal, A document analysis program according to claim 6, further comprising the following:

Citation Information

Patent Citations

  • Electron beam picture recorder

    JP1988005671A

  • Specification content inspection method, and specification content inspection system

    JP2009064339A

  • Requirement specification description support method

    JP2013105288A

  • Method of determining document similarity

    JP2015219799A