Article analysis device and article analysis method
The text analysis apparatus addresses ambiguity in requirement specifications by structurally analyzing and defining roles of phrases, enhancing the accuracy of requirement analysis and verification.
Patent Information
- Application Number
- JP2021149360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-14
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2041-09-14
AI Technical Summary
Existing systems face challenges in accurately interpreting and verifying requirements specifications due to ambiguous expressions in upstream development texts, making it difficult to perform accurate analysis and verification.
A text analysis apparatus and method that analyzes the structure of requirement texts, identifies semantic roles of phrases, and uses a dictionary database to extract and display definitions associated with these roles, reducing ambiguity and improving interpretation.
Enhances the accuracy of requirement specification analysis and verification by providing clear definitions for ambiguous phrases, improving the development process and ensuring compliance with specifications.
Smart Images

Figure 0007705322000001 
Figure 0007705322000002 
Figure 0007705322000003
Abstract
Description
Technical Field
[0001] The present invention relates to a text analysis apparatus and a text analysis method.
Background Art
[0002] In the development process of various systems in which software programs are processed, in order to ensure the quality of the system, verification is performed as to whether the system meets the required specifications. As an example of a technique for assisting such verification work, a technique has been proposed in which a machine-readable description of required specifications is created by combining selectable natural language text segments, and based on the description, a test for verifying whether the system meets the required specifications is automatically generated.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Thus, in order to determine how to verify whether the system meets the requirements specifications, it is desirable that the requirements specifications can be uniquely and unambiguously interpreted from the text in which the requirements specifications are described. However, the text describing the requirements specifications included in specifications and the like created in the upstream process of system development is more abstract than the software design documents and the like created in the downstream process, and ambiguous expressions are often used. In that regard, in the example of the prior art described above, it is proposed to generate the description of the requirements specifications itself in a machine-readable format. However, it is more desirable that the requirements specifications can be appropriately interpreted based on the description of normal requirements specifications created in the normal system development process, rather than such a special format description. Also, enabling appropriate interpretation of the requirements specifications based on such text is effective not only in the verification work of the system, but also in stages such as analyzing the requirements specifications of the system in the upstream process.
[0005] Therefore, in one aspect of the present invention, it is an object to enable appropriate interpretation of the requirements specifications of the system based on the requirement text in which the requirements specifications of the system are described.
Means for Solving the Problem
[0006] In one aspect of the present invention, input data including a requirement text in which requirements specifications for a system to be developed are described is read, the text structure analysis of the requirement text is performed, and the semantic role of the target phrase included in the requirement text in the requirements specifications is identified. Also, a dictionary database storing dictionary information in which definitions corresponding to the role are associated for each role for phrases related to the system is searched using the target phrase as a key, and the dictionary information corresponding to the target phrase is extracted from the dictionary database. Further, the definition corresponding to the role identified for the target phrase is extracted from the dictionary information. Then, output data in a form that can display the definition corresponding to the role of the target phrase together with the requirement text is generated and output.
Effect of the Invention
[0007] According to one aspect of the present invention, based on a requirement sentence describing the requirement specifications of a system, the requirement specifications of the system can be appropriately interpreted.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments for carrying out the present invention will be described in detail with reference to the accompanying drawings. <Technical Field to which the Present Embodiment is Applied> First, to facilitate understanding of the present embodiment, an example of the technical field (background) to which the present embodiment is applied will be described. The present embodiment relates to the development support of various systems in which a program is executed in a computer (including a microcomputer). As an example, in the development of an in-vehicle system, development is carried out based on a development process that applies a model called the V-shaped development model based on a process model such as A-SPICE (Automotive Software Process Improvement and Capability dEtermination). In the V-shaped development model, development is carried out in the order from the upstream process to the downstream process, that is, (1) requirement (specification) analysis of the system to be developed, (2) system architecture design, (3) software requirement analysis, (4) software architecture design, (5) software detailed design. Then, for each of the processes (1) to (5), verification operations of (6) software unit verification, (7) software integration test, (8) software qualification confirmation test, (9) system integration test, (10) system qualification confirmation test are carried out in the order from the downstream process to the upstream process to further improve the quality of the system.
[0010] Here, in the development process as described above, for example, in (1) which is an upstream process, a requirement specification document or the like in which the requirement specifications (specifications) for the system are described is referred to. The requirement sentences in which the requirement specifications are described, which are included in such requirement specification documents or the like, are often more abstract and ambiguous compared to detailed design documents or the like created in the software detailed design process of (5) for example. Therefore, for example, in the process of (1), it is difficult to correctly perform requirement analysis, and the part that depends on the skills of analysts or the like may become large. Also, when creating a test to confirm the system qualification in the process of (10) for verifying whether the process of (1) has been correctly performed, it may be difficult to determine by what kind of test it can be verified whether the system meets the requirement specifications.
[0011] In this embodiment, by analyzing the requirement statement created in the upstream process of the development process of such a system, the ambiguity of the requirement statement is reduced so that the requirement specifications can be clearly interpreted from the requirement statement. As a result, by improving the accuracy of the analysis work and verification work of the requirement specifications, the development support of the system is realized.
[0012] <First Embodiment> [Configuration of the Document Analysis Device] FIG. 1 is a diagram showing an example of a document analysis device according to the first embodiment. The document analysis device 1 is a computer and includes a control unit 10 that executes a program and a storage unit 20 that stores data used for the execution of the program. The control unit 10 includes a document structure analysis unit 11, a dictionary search unit 12, a definition extraction unit 13, and a document output unit 14, the functions of which are realized by the program being loaded from the storage unit 20 and executed. In addition to the data (not shown) of the program main body executed by the control unit 10, the storage unit 20 includes at least requirement statement input data 21, a dictionary database 22, a synonym phrase database 23, and requirement statement output data 24.
[0013] The document structure analysis unit 11 reads the requirement statement input data 21 from the storage unit 20 and performs document structure analysis processing on the requirement statement included in the requirement statement input data 21. The document structure analysis unit 11 identifies, by the document structure analysis processing, the semantic role in the requirement specifications described in the requirement statement for the phrases (which may include one word or a phrase formed by combining a plurality of words) included in the requirement statement. The phrases to be identified for such roles are, for example, nouns or noun phrases, but are not limited thereto. In this embodiment, such roles are classified into three roles: "condition", "observation target", and "expected operation". In the following description, the phrase to be identified for such a role in the requirement statement is referred to as a target phrase. The dictionary search unit 12 accesses the dictionary database 22 and extracts dictionary information corresponding to each target phrase obtained by the document structure analysis unit 11. The definition extraction unit 13 extracts a definition corresponding to the role of each target phrase analyzed by the sentence structure analysis unit 11 from the dictionary information extracted by the dictionary search unit 12. The sentence output unit 14 generates and outputs request sentence output data 24, which is data in a form that can display the definition corresponding to the role of the target phrase extracted by the definition extraction unit 13 together with the request sentence.
[0014] The request sentence input data 21 is data in which the request sentences included in the requirement specification in system development are stored in a format readable by a computer (for example, CSV (Comma Separated Values), etc.). FIG. 2 shows an example of the request sentence input data 21 in tabular form. The request sentence input data 21 includes, for example, a request ID for identifying each request sentence, the request sentence, and the chapter (section) in the requirement specification to which the request sentence belongs.
[0015] The dictionary database 22 is a database in which definitions (explanations of specific contents regarding the relevant phrases) are associated with phrases related to the system to be developed. FIG. 3 shows an example of the data structure of the dictionary database 22. As shown in FIG. 3, in the dictionary database 22, as dictionary information, definitions are associated with each phrase according to the semantic role in the requirement specifications as described above. In the present embodiment, as described above, such roles are classified into three roles: "condition", "observation target", and "expected operation". Note that the classification of roles is not limited to such examples and can be arbitrarily set according to the content and characteristics of the requirement sentences in the system. Also, it is not necessary to associate definitions of all three of these roles with all phrases, and for some phrases, only one of the three definitions may be associated. Further, as shown in FIG. 3, in the dictionary database 22, for each role, the probability that the relevant phrase corresponds to that role is associated. The dictionary information of the dictionary database 22 can be arbitrarily registered in all of the above-described system development processes. In particular, by registering dictionary information in the dictionary database 22 in the downstream process of the system development process, it is possible to provide more detailed definitions for the phrases. Such a dictionary database 22 can be created in a format such as JSON (JavaScript Object Notation), for example.
[0016] The synonym phrase database 23 is a database in which synonym phrases are associated with phrases related to the system to be developed. In the synonym phrase database 23, as shown in FIG. 4, for each phrase, one or more synonym phrases and the probability of substituting the relevant phrase with the synonym phrase are associated. The data of the synonym phrase database 23 can be registered in advance before starting system development, and can also be arbitrarily registered in all of the above-described system development processes.
[0017] The claim text output data 24 is data in a form that can display a definition according to the role of the target phrase together with the claim text. As an example, the claim text output data 24 can be composed of HTML (Hyper Text Markup Language) data including the claim text. FIG. 5 shows an example of a state where the claim text output data 24 is generated in the format of HTML data and the HTML data is output on a web browser. In the claim text output data 24, by incorporating JavaScript, CSS (Cascading Style Sheets), etc. into the HTML data, it is possible to implement to display the claim text on the web browser and dynamically display the definition corresponding to the role of each phrase as a comment.
[0018] [Article analysis process] FIG. 6 is a flowchart showing the article analysis process executed in the article analysis apparatus 1 according to the first embodiment. In step 1001 (denoted as S1001 in the figure. The same applies hereinafter), the sentence structure analysis unit 11 reads the claim text input data 21 from the storage unit 20. Then, the sentence structure analysis unit 11 performs a sentence structure analysis process on the claim text included in the claim text input data 21.
[0019] The sentence structure analysis process includes, for example, the following three processes. (1) Morphological analysis The sentence structure analysis unit 11 performs a morphological analysis that divides the claim text into morphemes, which are the smallest units having meaning in language, and discriminates the part of speech and inflection of each morpheme. Existing techniques can be used for morphological analysis. As specific examples, analysis tools such as MeCab and JUMAN are provided, and existing corpora, etc. can be used as the dictionaries used for the analysis process.
[0020] (2) Syntactic analysis Next, based on the results of morphological analysis, the sentence structure analysis unit 11 performs syntactic analysis (dependency analysis) to analyze the relationship between each word and phrase included in the claim sentence. Existing technologies can be used for syntactic analysis. For example, it is possible to utilize tools based on machine learning algorithms. As such tools, for example, there are recurrent neural network (RNN) models constructed using TensorFlow and SyntaxNet specialized for natural language understanding (NLU).
[0021] (3) Role analysis of words and phrases (semantic analysis) Based on the results of morphological analysis and syntactic analysis, the sentence structure analysis unit 11 identifies the role (whether it is a "condition", an "observation target", or an "expected operation") of the target word and phrase included in the claim sentence within the claim sentence. The classification of the role at this time shall be made to match the classification of the role set in the dictionary information of the dictionary database 22. Such role analysis processing can be realized, for example, by deep learning based on machine learning algorithms. As a specific example of existing technologies for realizing such deep learning, it is possible to use frameworks of neural networks described in Python and the like.
[0022] As an example of processing by deep learning, it is possible to prepare teacher data in which the roles of words used in a claim sentence are classified into "conditions", "observation targets", and "expected actions", perform learning processing using the teacher data, and generate a learning model. Further, as the input layer of deep learning, for example, the target phrase itself, the grammatical structure of the claim sentence, the chapter to which the claim sentence belongs in the requirement specification document, keywords included in the claim sentence, etc. can be included. Information indicating the grammatical structure of the claim sentence is, for example, information indicating what kind of words the words before and after the target phrase are. As a specific example, this information can be used for analysis such as that if there is a particle such as "in" immediately after the target phrase, the role is likely to be "condition". Also, the output layer can be information indicating any of the roles of the above-described phrases, that is, "condition", "observation target", and "expected action". Although not shown in FIG. 1, the teacher data and the data of the learning model used for the deep learning are also stored in the storage unit 20 of the document analysis device 1.
[0023] In step 1002, the dictionary search unit 12 accesses the dictionary database 22 and performs a process of searching for and extracting dictionary information corresponding to each target phrase included in the claim sentence. Details of the dictionary search process will be described later. As shown in FIG. 3, the dictionary information extracted from the dictionary database 22 by the dictionary search unit 12 includes definitions corresponding to all or part of the roles of the target phrases in the claim sentence, that is, "conditions", "observation targets", and "expected actions", and information indicating the probability of each.
[0024] In step 1003, the definition extraction unit 13 extracts, from the dictionary information, a definition corresponding to the role of the target phrase identified by the sentence structure analysis unit 11 for each of the target phrases included in the claim sentence from which the dictionary information has been extracted from the dictionary database 22.
[0025] In step 1004, the sentence output unit 14 generates and outputs request sentence output data 24 that can indicate the definition of a target phrase included in the request sentence together with the request sentence for the target phrase for which a definition corresponding to the role has been extracted. As described above, as an example, as shown in FIG. 5, the sentence output unit 14 can display the request sentence on a web browser and dynamically display the definition corresponding to the role of each target phrase as a comment. Thereby, the user can refer to the definition corresponding to the role of each target phrase together with the request sentence. Also at this time, as also shown in FIG. 5, the sentence output unit 14 can also output the probability set in the dictionary database 22 for the definition corresponding to the role of the phrase.
[0026] FIG. 7 is a flowchart showing the detailed content of the dictionary search process in step 1002 of the sentence analysis process shown in FIG. 6. In step 1011, the dictionary search unit 12 accesses the dictionary database 22 and performs a search using the target phrase included in the request sentence as a key. In step 1012, the dictionary search unit 12 determines whether the search hits, that is, whether dictionary information corresponding to the target phrase exists in the dictionary database 22. If the dictionary information exists, the process proceeds to step 1013 (Yes), and if it does not exist, the process proceeds to step 1014 (No). In step 1013, the dictionary search unit 12 extracts the dictionary information corresponding to the target phrase from the dictionary database 22.
[0027] In step 1014, the dictionary search unit 12 performs a fuzzy search on the target phrase in the dictionary database 22. Specifically, the dictionary search unit 12 searches the synonym phrase database 23 using the target phrase as a key. When the search hits, that is, when there is data of a synonym phrase corresponding to the target phrase in the synonym phrase database 23, the dictionary search unit 12 extracts from the data of the synonym phrase the synonym phrase and the probability of substituting the target phrase with the synonym phrase. In this fuzzy search, if there is a phrase registered in the synonym phrase database 23 corresponding to a part of the phrase, the substitution of that part of the phrase with a synonym phrase may be performed. Also, if the probability value is less than or equal to a predetermined threshold, the extraction of the synonym phrase may not be performed. Then, the dictionary search unit 12 performs a re-search of the dictionary database 22 using the extracted synonym phrase as a key. In step 1015, the dictionary search unit 12 determines whether dictionary information corresponding to the target phrase exists in the dictionary database as a result of the search. If the dictionary information exists, it proceeds to step 1013 (Yes), and if it does not exist, it proceeds to step 1016 (No).
[0028] In step 1016, the dictionary search unit 12 performs sentence replacement processing on the requested sentence. Specifically, the dictionary search unit 12 estimates the dependency relationship between phrases using the existing syntax analysis technology as described above, and changes the word order so that the meaning of the sentence remains unchanged. For example, assume that the granularity of a certain phrase is larger than the phrases in the dictionary information included in the dictionary database 22 (i.e., it is a higher-level concept than the phrases in the dictionary information included in the dictionary database 22). In this case, the modifying phrase in the dependency relationship that modifies the phrase may exist at another position in the requested sentence that is not consecutive with the phrase. In this case, by replacing the modifying phrase at the position consecutive with the phrase to form one combined phrase, the phrase will have a more specific meaning. At this time, the particles before and after the phrase whose position has been changed may be changed as necessary. Then, by combining a plurality of phrases into one phrase in this way, or by combining each of the phrases at consecutive positions after being converted into synonymous phrases into one phrase, the phrase may match the phrase included in the dictionary database 22.
[0029] After performing the sentence replacement processing in step 1016, the dictionary search unit 12 proceeds to step 1011 again, and performs a search of the dictionary database 22 again in the state where the sentence replacement processing has been performed. Note that the number of times of performing the processing in steps 14 to 16 when the search of the dictionary database 22 fails may be counted, and when the number of times exceeds a predetermined threshold, the dictionary information search processing may be terminated.
[0030] [Data Specific Example of Sentence Analysis Processing] FIG. 8 is an explanatory diagram showing a data specific example of the sentence analysis processing described with reference to FIG. 6. Here, the requested sentence of the requested ID "ID-001" of the requested sentence input data 21 shown in FIG. 2, "The vehicle speed in the evacuation driving mode is limited to 50 km / h or less" is the processing target. As for the claim sentence, as a result of the morphological analysis and syntactic analysis in the sentence structure analysis section 11 of the sentence structure analysis process (S1001), the claim sentence can be split, for example, into "The vehicle speed in the evacuation driving mode is limited to 50 km / h or less." Among these, the part "in the evacuation driving mode" can be analyzed as a modifier (M) due to the description "in". The part "The vehicle speed is" can be analyzed as a subject (S) due to the auxiliary word "is". The part "to 50 km / h or less" can be analyzed as a complement (C) due to the auxiliary word "to". Also, the part "limit" can be analyzed as a predicate (V) because its part of speech is a verb, etc.
[0031] Based on the analysis result, the roles of each phrase are further identified. For example, "the evacuation driving mode" can be identified as having the role of "condition" in the claim sentence due to being included in the modifier and the description "in". "The vehicle speed" can be identified as having the role of "observation target" in the claim sentence because it is included in the subject. "50 km / h or less" and "limit" can be identified as having the role of "expected action" in the claim sentence because they are included in the complement and the predicate, etc.
[0032] Next, for each phrase, the dictionary search section 12 searches for dictionary information. Specifically, it accesses the dictionary database 22, searches the dictionary database 22 using each phrase as a key, and extracts the dictionary information (S1002, S1011 to S1016). In this specific example of the data, it is assumed that dictionary information is extracted for each of "the evacuation driving mode" and "the vehicle speed" from the dictionary database 22 shown in FIG. 3. Each dictionary information includes the classifications of the roles of the phrases in this embodiment, namely "condition", "observation target", and "expected action", and the definitions and probabilities associated with each of these roles.
[0033] Then, the definition extraction unit 13 extracts the definitions and probabilities according to their respective roles from the dictionary information extracted for each of "evacuation driving mode" and "vehicle speed". In this specific example of data, since the role of "evacuation driving mode" is identified as "condition", from the dictionary information, the definition "When both diagnosis ○○ and ×× are NG, the flag △△ is set and the evacuation driving mode is entered" and the probability "85 percent" are extracted. Also, for "vehicle speed", since its role is identified as "observation target", from the dictionary information, the definition "RAM value: ○○〇 (operation cycle xx ms), received value: ××× (CAN), raw value measurement method: △△△" and the probability "75 percent" are extracted.
[0034] Furthermore, the sentence output unit 14 generates and outputs a required sentence output data 24 that can display the definitions extracted for these "evacuation driving mode" and "vehicle speed" together with the required sentence. As described above, FIG. 5 shows an output example of the required sentence output data 24. In this output example, on the screen displayed on the computer display, the required sentence "The vehicle speed in the evacuation driving mode is limited to 50 km / h or less" is displayed. And for the phrases "evacuation driving mode" and "vehicle speed" for which the definitions according to the roles are extracted in the required sentence, underlines are displayed. When the user overlays the mouse pointer on these phrases, the role of the phrase, the definition according to the role, and the probability are dynamically displayed as comments.
[0035] Next, regarding the ambiguity search (S1014) and sentence replacement (S1016) in the dictionary search process, which was particularly described with reference to FIG. 7 in the text analysis process, a specific data example shown in FIG. 9 will be referred to for explanation. Here, the sentence to be processed is "The inverter must be switched off when it receives a shutdown request", which is the request sentence of the request ID "ID-002" in the request sentence input data 21 shown in FIG. 2. On the other hand, as shown in FIG. 3, the dictionary database 22 has the phrase "power module shutdown" registered. Also, as shown in FIG. 4, in the synonym phrase database 23, the synonym phrases and probabilities for the phrases "inverter" and "switch off" are registered.
[0036] Here, when an ambiguity search is performed on the said request sentence, synonym phrases can be extracted for each of the phrases "inverter" and "switch off" in the synonym phrase database 23, but the phrases corresponding to each of the synonym phrases are not registered in the dictionary database 22. Therefore, when a sentence replacement process is performed on the said request sentence, through the analysis of its syntactic relationship, etc., the said request sentence can be changed to "The inverter switch-off must be done when it receives a shutdown request". Then, when an ambiguity search is performed again on the sentence after the said sentence replacement, for the phrase "inverter switch-off", the synonym phrase "power module" for "inverter" and the synonym phrase "shutdown" for "switch off" can be extracted respectively. In this way, in the said request sentence, the phrase "power module shutdown" combined from these can be recognized. And it becomes possible to extract the definition of the phrase "power module shutdown" from the dictionary database 22.
[0037] In this case, the probability of substituting a synonymous phrase for "power module shutdown" can be set to "35%", which is obtained by multiplying "50%" of "power module" and "70%" of "shutdown" in the synonymous phrase database 23 shown in FIG. 6. Further, when such a fuzzy search is performed, in the output example of the required sentence output data 24 shown in FIG. 5, the probability of displaying the definition according to the role of the phrase may be, for example, a value obtained by multiplying the value of the probability of the definition according to the role and the value of the probability of substituting a synonymous phrase. Also, when the sentence replacement process and the fuzzy search process are performed in this way, the sentence output unit 14 may generate the required sentence output data 24 so as to display both the original text of the required sentence and the sentence converted in the sentence replacement process and the fuzzy search.
[0038] [Effects etc. according to the first embodiment] As described above, in this embodiment, in a sentence analysis device that analyzes a required sentence in which the required specifications for the system to be developed are described, the sentence structure of the required sentence is analyzed, and the role of the target phrase included in the required sentence is identified. Further, a dictionary database 22 in which dictionary information in which definitions according to roles are associated for each role is stored for phrases related to the system is searched using the target phrase as a key, and the dictionary information corresponding to the target phrase is extracted from the dictionary database 22. Then, the definition associated with the role identified for the target phrase is extracted from the dictionary information, and the required sentence is output in a manner that allows the definition according to the role of the target phrase to be displayed. As a result, the user involved in system development can refer to the definition according to the role of each phrase together with the required sentence.
[0039] Here, in particular, the ability to refer to definitions according to the roles of each phrase has the following effects. That is, even for the same phrase, for example, the meaning (object) indicated by the phrase may differ depending on the context, etc. in each claim statement. Therefore, even if a definition uniquely set for the phrase can be referred to, the definition is not necessarily accurate in each claim statement. In this regard, in the present embodiment, in the dictionary database 22, definitions according to the roles are associated with each phrase for each role. Then, for the target phrase included in the claim statement, a definition according to the role identified for the target phrase is extracted. As a result, the user can refer to an accurate definition according to the content of the requirement specification.
[0040] Therefore, in the present embodiment, even for a claim statement included in a requirement specification or the like, which is more abstract than a software design document or the like created in the downstream process of system development and often uses ambiguous expressions, the ambiguity can be reduced. As a result, the user can appropriately interpret the requirement specification based on the claim statement. Thereby, the accuracy of the requirement specification analysis work and verification work is improved, and the development support of the system is realized.
[0041] For example, in the case of the V-shaped development model described above, through the downstream process of the development process, a large number of dictionary information including detailed definitions according to the roles are registered in the dictionary database 22 for phrases related to the system to be developed. Therefore, in particular, in the subsequent verification process of requirement analysis (system qualification confirmation test), the hit rate in the search of the dictionary database 22 is improved. Therefore, it becomes possible to extract definitions according to the roles of more target phrases for the claim statement, and the accuracy of the verification work of requirement analysis is improved.
[0042] Here, as an example of the system to be developed to which the present embodiment is applied, consider the case of a mechatronics system that senses and controls up to the dynamics of the control target, such as a control device for an automotive powertrain. In this case, the requirements specifications that the control device should meet include the behavior of the powertrain, such as the engine and transmission of the vehicle being controlled. For example, regarding the driver's accelerator operation, it is required to obtain a sensory evaluation of the ride quality, such as whether the expected acceleration can be obtained when adjusting the throttle opening and fuel injection amount. The requirements text that describes such requirements specifications inevitably contains ambiguity. On the other hand, in the actual development process, it will go through downstream processes such as the design and detailed design of the system architecture created by considering specifically which sensors and actuators to control when and how. However, because the system is complex, the relevance between the design content of the individual components of the system created in these downstream processes and the original requirements specifications tends to become weak. Therefore, there has been a difficult situation in determining how to specifically verify whether the original requirements specifications that should be verified are satisfied. In such a situation, according to the present embodiment, for the target phrases included in the requirements specifications, specific definitions set in the downstream process can be referred to according to their roles. As a result, it becomes possible to easily grasp the relevance between the requirements specifications and the components of the individual systems, such as which functions of which components in the detailed design and the like should be realized to meet the requirements specifications.
[0043] Also, in the present embodiment, the roles are classified into "conditions", "observation targets", and "expected operations", and definitions corresponding to these roles are extracted. As a result, it becomes possible to clarify what conditions the requirements specifications described in the requirements text are based on, what the observation target is, and furthermore, how the observation target should operate. Such a classification of roles is particularly effective for interpreting the requirements specifications for a system including control devices.
[0044] Also, in the present embodiment, when the dictionary information associated with the target phrase does not exist in the dictionary database 22, a synonymous phrase associated with the target phrase is extracted from the synonymous phrase database 23. Then, the dictionary information associated with the synonymous phrase is extracted from the dictionary database 22, and from the dictionary information, a definition associated with the role identified for the target phrase is extracted. Then, the definition of the target phrase is displayed, and the plausibility of the synonymous phrase associated with the target phrase is also displayed. As a result, even if the target phrase is not registered in the dictionary database 22, if its synonymous phrase is registered in the dictionary database 22, it becomes possible to extract the definition associated with the role of the target phrase. Therefore, it is possible to improve the hit rate in the search of the dictionary database 22 as a result, and to further reduce the ambiguity of the claim text.
[0045] Furthermore, in the present embodiment, text replacement for the claim text, that is, a process of estimating the syntactic relationship between phrases and changing the order of the phrases so that the meaning of the text does not change is performed. As a result, as described above, for example, two phrases at different positions are changed to consecutive positions, and a case may occur where the phrases are concatenated and match one phrase registered in the dictionary database 22. Therefore, it is possible to improve the hit rate in the search of the dictionary database 22 as a result, and to further reduce the ambiguity of the claim text.
[0046] When registering the dictionary information in the dictionary database 22, a sub-structure may be further set according to the content for each role definition. As a specific example, in the development of an in-vehicle system, for the term "vehicle speed", assume that the calculation locations are "sensor value, backup sensor value, internal calculation value (motor rotation speed × gear ratio × tire diameter), CAN communication reception value". In this case, for example, each of these calculation locations may be set as the sub-structure of the definition corresponding to the role of "observation target". And when the sentence output unit 14 generates the required sentence output data 24, for example, all the contents of these sub-structures may be made referenceable in the display of the definition of the term. Also, the user may be allowed to select any of the information of these sub-structures.
[0047] <Second Embodiment> [Outline of the Second Embodiment (Differences from the First Embodiment)] Next, the second embodiment will be described. In the first embodiment, there is one dictionary database, and it was assumed that the dictionary information was mainly registered during the development process. In this case, among the development processes of the system as described above, in the verification stage for the upstream process (for example, the process (10) in the case of the V-shaped development model described above), the dictionary information is sufficiently registered through the downstream process. However, in the upstream process itself before entering the detailed design stage (for example, the process (1)), the dictionary information is not yet sufficiently registered.
[0048] Therefore, in the second embodiment, in the storage unit of the sentence analysis device, the dictionary databases already created in the systems developed in the past are accumulated, and the dictionary databases used in the systems with high affinity to the system to be developed are diverted. As a result, in the second embodiment, a large number of dictionary information including detailed definitions are registered in the dictionary database from the start point of system development. Therefore, even in the upstream process before entering the detailed design stage, for the target terms included in the required sentences, detailed definitions corresponding to their roles can be made referenceable, and development support can be realized.
[0049] [Configuration of the Document Analysis Device] FIG. 10 shows the configuration of the document analysis device 1 in the second embodiment. Different from the first embodiment, the document analysis device 1 in the second embodiment further includes a dictionary selection unit 15 in the control unit 10. Further, the storage unit 20 further includes request document past data 25. The request document past data 25 can be stored in the number corresponding to the number of other systems that have been developed in the past. Also, in the second embodiment, a plurality of dictionary databases 22 and thesaurus databases 23 can also be stored. Since the other components are the same as those in the first embodiment, the description thereof is omitted.
[0050] Based on the request document input data 21 and the request document past data 25 of the systems that have been developed in the past, the dictionary selection unit 15 identifies a system with high affinity between the system to be developed this time and the developed systems. Then, the dictionary database 22 and thesaurus database 23 used in the system are selected. The request document past data 25 is, for example, one or more data having the same structure as the request document input data 21, and the request documents of one or more systems that have been developed in the past are stored data for each system.
[0051] The dictionary database 22 has the same structure as that in the first embodiment, but a plurality of them may exist in the second embodiment. And each dictionary database 22 is associated with the request document past data 25 in which the request specifications of other systems where the dictionary database 22 was used in the past are stored. Each dictionary database 22 stores data that has already been created in the development process of other systems.
[0052] The synonym sentence database 23 also has the same structure as that in the first embodiment, but there may be a plurality of them in the second embodiment. Similar to the dictionary database 22, each synonym sentence database 23 is associated with requirement sentence past data 25 in which the requirement specifications of other systems where the synonym sentence database 23 has been used in the past are stored. Each synonym sentence database 23 stores data that has already been created in the development process of other systems. Note that the synonym sentence database 23 does not necessarily have to be a selection target of the dictionary selection unit 15. One synonym sentence database 23 may be used in all systems, and synonyms may be further registered and accumulated as needed. Alternatively, the data of the synonym sentences in the synonym sentence database 23 may be commonly used in all systems, and only the certainty of each synonym sentence may be set to a value different depending on the characteristics of the system. In this case, the dictionary selection unit 15 will select data with a high degree of affinity for the system to be developed this time.
[0053] [Text analysis process] FIG. 10 is a flowchart showing the text analysis process executed in the text analysis apparatus 1 in the second embodiment. In step 2001, the dictionary selection unit 15 identifies a system with high affinity for the system to be developed this time based on the requested text input data 21 and the requested text past data 25 of other systems that have been developed in the past. For example, the dictionary selection unit 15 can identify a system with high affinity in the following manner. First, the dictionary selection unit 15 reads out each of the requested text input data 21 and the requested text past data 25. Then, semantic analysis (latent semantic analysis) can be performed on the requested text included in the requested text input data 21 and the requested text included in the requested text past data 25. Further, the analysis result of the requested text included in the requested text input data 21 is compared with the analysis result of the requested text included in the requested text past data 25. The dictionary selection unit 15 performs this process for all the requested text past data 25 stored in the storage unit 20, and based on each comparison result, identifies one of the requested text past data 25.
[0054] More specifically, for example, as an example of the semantic analysis result of the requested text, in a vector space based on the granularity of nouns and noun phrases, a numerical value representing the usage frequency of each sentence can be obtained. The dictionary selection unit 15 calculates the logical sum of the analysis result of the requested text included in the requested text input data 21 and the analysis result of the requested text included in each of the requested text past data 25, and can perform a gap analysis between the two based on the logical sum. Then, by identifying the requested text past data 25 with the smallest gap, that is, the highest similarity, to the requested text input data 21, a system with high affinity for the system to be developed this time can be identified.
[0055] Then, the dictionary selection unit 15 selects the dictionary database 22 associated with the identified requested text past data 25, that is, the dictionary database 22 used in a system with high affinity for the system to be developed this time. Similarly, the dictionary selection unit 15 can also select the synonym sentence database 23 used in a system with high affinity for the system to be developed this time.
[0056] Steps S2002 to S2005 are the same as S1001 to S1004 in the text analysis process of the first embodiment, so the description is omitted. Also, in the second embodiment, S1011 to S1016 of the detailed content of the dictionary search process shown in FIG. 7 are similarly executed.
[0057] [Effects etc. according to the second embodiment] According to the second embodiment, regarding the dictionary database used in the text analysis process, it is possible to reuse another dictionary database that has already been created in a system similar to the system to be developed and that has been developed in the past. As a result, in the second embodiment, from the start of system development, a dictionary database with a large number of registered dictionary information including detailed definitions can be used. Therefore, not only in the verification stage of system development, but also in the upstream process before entering the detailed design stage, it becomes possible to refer to detailed definitions according to the role of the target phrases included in the required text, and it becomes possible to realize system development support.
[0058] Note that when identifying a system similar to the system to be developed, a predetermined threshold for determining the similarity may be set. And when the similarity is lower than the predetermined threshold (which is synonymous with the result of the gap analysis being higher than the predetermined threshold), instead of reusing the existing dictionary database 22, a new dictionary database 22 may be generated. The same applies to the synonym phrase database 23. Also, when there is only one piece of required text past data 25 stored in the storage unit 20 in advance, or the dictionary database 22 and the synonym phrase database 23 used in the past system, there is no need to perform the process of identifying a similar system. Also in this case, as described above, when the similarity is lower than the predetermined threshold, the dictionary database 22 and the synonym phrase database 23 may not be reused.
[0059] [Hardware configuration] FIG. 12 shows an example of the hardware configuration of a computer that constitutes the text analysis apparatus 1 in each of the above embodiments. This computer includes a processor 910, a memory 920, a storage 930, a removable storage medium drive 940, an input / output device 950, and a communication interface 960. The processor 910 includes a control unit, an arithmetic unit, an instruction decoder, etc. The execution unit executes arithmetic and logical operations using the arithmetic unit according to the control signal output from the control unit in accordance with the program instructions decoded by the instruction decoder. Such a processor 910 includes a control register in which various information used for control is stored, a cache that can temporarily store the contents of the memory 920 and the like that have already been accessed, and the like. Note that the processor 910 may have a configuration in which a plurality of CPU (Central Processing Unit) cores are provided. When a program is executed by the processor 910, the control unit 10 described above is configured.
[0060] The memory 920 is a storage device such as a RAM (Random Access Memory), and is a main memory in which the program executed by the processor 910 is loaded and the data used for the processing of the processor 910 is stored. The storage 930 is a storage device such as an HDD (Hard Disk Drive) or a flash memory, and stores programs and various data. The storage unit 20 described above corresponds to, for example, the storage 930. The removable storage medium drive 940 is a device that reads data and programs stored in the removable storage medium 970. The removable storage medium 970 is, for example, a magnetic disk, an optical disk, a magneto-optical disk, or a flash memory. Note that the processor 910 executes the programs stored in the storage 930 and the removable storage medium 970 while cooperating with the memory 920 and the storage 930. Note that the programs executed by the processor 910 and the data to be accessed may be stored in another device that can communicate with the computer.
[0061] The input / output device 950 is input means and output means such as a touch panel, a keyboard, etc. and a display, etc., which receives operation commands by user operations, etc., while outputting the processing results by the information processing device. The communication interface 960 enables data communication with the outside. Each component of the computer described above is connected by a bus 980.
[0062] <Others> The embodiments of the present invention described above are only a part of the embodiments conceivable within the technical scope of the present invention and do not limit the technical scope of the present invention. Also, the functional configuration and physical configuration in each embodiment are not limited to the above-described aspects. For example, each function and physical resource may be integrated and implemented, or conversely, may be further distributed and implemented, and furthermore, it is also possible to add, delete, replace, etc. a part of the configuration with other configurations.
Explanation of Reference Numerals
[0063] 1... Text analysis device, 11... Text structure analysis unit, 12... Dictionary search unit, 13... Definition extraction unit, 14... Text output unit, 15... Dictionary selection unit, 21... Request text input data, 22... Dictionary database, 23... Synonymous phrase database, 24... Request text output data, 25... Request text past data
Claims
1. A sentence analysis device that analyzes a requirement sentence in which requirement specifications for a system to be developed are described, a sentence structure analysis unit that reads input data including the requirement sentence, performs sentence structure analysis of the requirement sentence, and identifies the semantic role of a target phrase included in the requirement sentence in the requirement specifications; a dictionary search unit that searches a dictionary database storing dictionary information in which definitions corresponding to the roles are associated with phrases related to the system for each role, using the target phrase as a key, and extracts the dictionary information corresponding to the target phrase from the dictionary database; a definition extraction unit that extracts the definition corresponding to the role identified for the target phrase from the dictionary information; and a sentence output unit that generates and outputs output data in a form capable of displaying the definition corresponding to the role of the target phrase together with the requirement sentence. A sentence analysis device comprising the above components.
2. The sentence analysis device according to claim 1, wherein the role is any one of a condition, an observation target, and an expected operation in the requirement specifications.
3. When the dictionary information corresponding to the target phrase does not exist in the dictionary database, the dictionary search unit searches a synonym database in which synonyms of the phrase and the probability of replacing the phrase with the synonym are associated with phrases related to the system, using the target phrase as a key, extracts the synonym corresponding to the target phrase from the synonym database, searches the dictionary database again using the synonym as a key, and extracts the dictionary information corresponding to the synonym from the dictionary database. The definition extraction unit extracts the definition corresponding to the role identified for the target phrase from the dictionary information corresponding to the synonym. The sentence output unit according to claim 1 or 2, wherein the sentence output unit displays the definition extracted for the synonym and also displays the probability.
4. Read the input data including the claim text and each of one or more pieces of past data each including a claim text in which claim specifications for one or more other systems developed in the past are described for each of the other systems, perform semantic analysis to analyze the meaning of the claim text included in the input data and the claim text included in each of the past data, identify one piece of past data having a high degree of similarity to the input data among the past data based on the result of comparing the analysis result of the claim text included in the input data and the analysis result of the claim text included in each of the past data, and further include a dictionary selection unit that selects a dictionary database that is a dictionary database created in the development of the other system and is associated with the identified one piece of past data. The dictionary search unit searches the dictionary database selected by the dictionary selection unit using the target phrase as a key, and extracts the dictionary information corresponding to the target phrase from the dictionary database. The document analysis apparatus according to any one of claims 1 to 3.
5. A document analysis method for analyzing by a computer a claim text in which claim specifications for a system to be developed are described, The computer reads input data including the claim text, performs syntactic analysis of the claim text, and identifies the semantic role of the target phrase included in the claim text in the claim specifications; The computer searches a dictionary database storing dictionary information in which definitions corresponding to the role are associated with each role for phrases related to the system using the target phrase as a key, and extracts the dictionary information corresponding to the target phrase from the dictionary database; The computer extracts the definition corresponding to the role identified for the target phrase from the dictionary information; The computer generates and outputs output data in a form capable of displaying the definition corresponding to the role of the target phrase together with the claim text; A document analysis method including the above.
Citation Information
Patent Citations
Text base retrieving method
JP1990253474A
Natural language processor
JP1994019965A
Method and device for classifying document and recording medium with program recorded thereon
JP2000339310A
Method for creating requirement description for embedded system
JP2008171391A
Synonym extraction system, method and program
JP2013020439A