Classification support system, classification support device, classification support method, and program
The classification support system uses machine learning to efficiently classify RFP requirements by converting text into vectors, improving the accuracy and efficiency of functional and non-functional requirement extraction.
Patent Information
- Application Number
- JP2019003266
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2019-01-11
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2039-01-11
AI Technical Summary
The challenge lies in efficiently classifying functional and non-functional requirements from semi-structured documents like RFPs without compromising quality, as existing methods rely heavily on experienced engineers and are inefficient due to the complexity and scale of these documents.
A classification support system utilizing machine learning to classify requirements by converting text into word vectors and applying a trained model to identify and output classification results, enhancing efficiency and accuracy.
This approach improves the classification process by automating the extraction and classification of requirements, reducing variability and enhancing the quality of work, thus minimizing errors and delays.
Smart Images

Figure 0007757023000001 
Figure 0007757023000002 
Figure 0007757023000003
Abstract
Description
[Technical Field]
[0001] The embodiments of the present invention relate to a classification support system and a classification support device. , minutes This document relates to support methods and programs. [Background technology]
[0002] In the upstream process of system development, the system development contractor (hereinafter simply referred to as the "contractor") proposes a system with the functions required by the system development contractor (hereinafter simply referred to as the "client") based on the RFP (Request For Proposal; request for proposal or procurement specification) presented by the client. The RFP lists the system's functional requirements, non-functional requirements, and constraints. Functional requirements are requirements related to functions that must be included in the system. Non-functional requirements are requirements other than functional requirements, such as requirements related to system performance, usability, reliability, operability, availability, safety, migration, and scalability. Constraints are matters related to restrictions during system development and operation, such as contract terms, contractor qualifications, development organization, project management method, development environment, and operating environment.
[0003] In order to propose a system that meets the client's requirements, the contractor must comprehensively extract information on the functional requirements, non-functional requirements, and constraints described in the RFP. This is because omissions in the extraction of functional requirements, non-functional requirements, and constraints can lead to rework in later processes, which can cause problems such as delivery delays and cost overruns.
[0004] Typically, RFPs are semi-structured documents written in natural language. RFPs are created based on specific guidelines. However, these guidelines generally specify the table of contents and quality characteristics, but do not specify the classification of functional requirements, non-functional requirements, and constraints, and their correspondence with specific terms. Therefore, it is difficult to classify functional requirements, non-functional requirements, and constraints based on the guidelines. Furthermore, depending on the scale of the system, RFPs can be huge documents, for example, hundreds of pages or more. Therefore, comprehensively extracting functional requirements, non-functional requirements, and constraints from a massive RFP and accurately and efficiently classifying the extracted information has traditionally required the know-how of experienced engineers. As a result, there can be significant differences in work efficiency and work quality between experienced and inexperienced engineers who extract and classify functional requirements, non-functional requirements, and constraints from RFPs. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-127397 [Patent Document 2] Japanese Patent Application Laid-Open No. 2008-250760 [Patent Document 3] Japanese Patent Application Laid-Open No. 2012-243194 [Patent Document 4] Japanese Patent Application Publication No. 2014-203228 [Patent Document 5] Japanese Patent Application Publication No. 2017-224126 Summary of the Invention [Problem to be solved by the invention]
[0006] The problem that the present invention aims to solve is to provide a classification support system, classification support device, learning device, classification support method, and program that can improve the efficiency of the classification work of information contained in semi-structured documents without compromising the quality of the work. [Means for solving the problem]
[0007] The classification support system of the embodiment includes a classification learning unit, a classification unit, and an analysis unit; The classification learning unit has: It contains at least text that can be classified into functional requirements, which indicate requirements related to functions that must be installed in the system, and non-functional requirements, which indicate requirements other than the functional requirements. Learning information based on character strings contained in semi-structured learning documents and the content of the character strings The aforementioned The classification unit acquires learning data associated with classification information that identifies a classification, inputs learning input information based on the learning information and the classification information, performs machine learning, and stores a trained model that indicates a trained learning model in a storage unit. The requirements include at least text that can be classified into functional requirements and non-functional requirements. Classification information based on character strings contained in semi-structured documents is obtained, and classification input information based on the classification information is input to the trained model to identify the content of the character strings. The classification into the functional requirements and the non-functional requirements Regarding classification results information indicating a probability value for each of the classifications; Get classification result related information. The analysis unit generates classification result information indicating the classification result based on the classification result related information and a predetermined threshold value. The classification result output part is record Outputs similar result information. The analysis unit updates the value of the threshold based on the result of analyzing the classification result related information. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 10 is a diagram showing an example of a plurality of predefined categories and requirement specification texts classified into each category. [Figure 2] 1 is an overall configuration diagram of a classification support system 1. [Figure 3] FIG. 1 is a block diagram showing the functional configuration of a learning data creation device 10. [Figure 4] FIG. 2 is a diagram showing the table configuration of a requirement specification classification table T1. [Figure 5] FIG. 2 is a diagram showing the functional configuration of a learning device 20. [Figure 6] FIG. 2 is a diagram showing the functional configuration of a classification support device 30. [Figure 7] 4 is a flowchart showing the operation of the learning data creation device 10. [Figure 8] 4 is a flowchart showing the operation of the learning device 20 in generating a word vector model M1. [Figure 9] 10 is a flowchart showing the operation of the learning device 20 in generating a requirement specification classification model M2. [Figure 10] 4 is a flowchart showing the operation of the classification support device 30. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, a classification support system, a classification support device, a learning device, a classification support method, and a program according to embodiments will be described with reference to the drawings.
[0010] The classification support system 1 of the embodiment described below is a system for supporting the classification of character strings extracted from semi-structured documents into one of multiple predefined categories. Examples of semi-structured documents include requirement specification documents for system development, application documents for government offices, and questionnaire results including free-form comment fields.
[0011] The classification support system 1 described below is a system that supports the classification of strings extracted from requirement specification documents in system development (hereinafter referred to as "requirement specification text") into one of several predefined categories: functional specifications, non-functional specifications, or constraints.
[0012] Specific examples of the above-mentioned predefined divisions will be described below. FIG. 1 is a diagram showing an example of multiple predefined categories and requirements specification text classified into each category. As shown in FIG. 1, requirements specification text extracted from a requirements specification document is classified into one of the following categories (major categories): functional requirements, non-functional requirements, or constraints. Furthermore, requirements specification text classified into non-functional requirements and constraints is further classified into one of multiple categories (minor categories). For example, the major category "non-functional requirements" is further classified into one of the minor categories, such as "performance," "reliability," or "operability." Furthermore, for example, the major category "constraints" is further classified into one of the minor categories, such as "contractor qualifications" or "project management method."
[0013] Functional requirements are requirements related to functions that a system must have. As shown in Figure 1, requirement specification text classified as "functional requirements" is a string that expresses a requirement, such as "It must be able to link with the XX system and obtain XX information." Non-functional requirements are requirements other than functional requirements. As shown in Figure 1, requirement specification text classified as "performance" under "non-functional requirements" is a string that expresses a requirement, such as "It must be able to process XX items of online data per hour." Constraints are matters related to restrictions imposed during system development and operation. As shown in Figure 1, requirement specification text classified as "contractor qualifications" under "constraints" is a string that expresses a matter such as "The contractor must ensure the availability of personnel with the necessary skills and experience throughout the entire period of this work to ensure the work is carried out reliably."
[0014] The overall configuration of the classification support system 1 according to the embodiment will be described below. Fig. 2 is an overall configuration diagram of the classification support system 1. As shown in Fig. 1, the classification support system 1 includes a learning data creation device 10, a learning device 20, and a classification support device 30. The learning data creation device 10, the learning device 20, and the classification support device 30 each include an information processing device such as a personal computer.
[0015] The learning data creation device 10 is a device for creating learning data (teacher data) used for machine learning (hereinafter simply referred to as "learning") performed by the learning device 20. A learning requirement specification text is input to the learning data creation device 10. The input learning requirement specification text is classified, for example, by an experienced engineer operating the learning data creation device 10, and associated with the classification results. The learning data creation device 10 outputs data (hereinafter referred to as "learning data") in which the learning requirement specification text is associated with the classification results to the learning device 20.
[0016] The learning device 20 is a device that generates a trained model for classifying requirement specification text. The learning device 20 performs preprocessing, such as natural language processing, and conversion to word vectors (word vectorization) on the training requirement specification text included in the training data output from the training data creation device 10. The conversion to word vectors is performed using a word vector model generated by machine learning.
[0017] The word vector model is generated in advance by machine learning using, for example, a corpus obtained from the Internet or training requirement specification text as training data. A corpus is a linguistic resource consisting of information such as character strings contained in newspapers, magazines, books, etc., collected in large quantities, or transcribed spoken language, which has been processed so that it can be searched and analyzed by a computer. In this embodiment, the word vector model is generated by machine learning using the corpus as training data.
[0018] The learning device 20 performs machine learning using data in which the learning requirement text converted into word vectors and the classification results are associated as training data, to generate a requirement classification model, which is a trained model for classifying the requirement text. The learning device 20 outputs the generated trained models (i.e., the word vector model and the requirement classification model) to the classification support device 30.
[0019] The classification support device 30 acquires requirement specification text to be classified. The classification support device 30 performs preprocessing, such as natural language processing, on the acquired requirement specification text and converts it into word vectors using a word vector model. The classification support device 30 inputs the requirement specification text converted into word vectors into a requirement specification classification model to obtain information on the classification results of the requirement specification text. The classification support device 30 then analyzes the information on the classification results of the requirement specification text and outputs information indicating the classification results.
[0020] In this embodiment, the training data creation device 10, the learning device 20, and the classification support device 30 are each separate devices, but this is not limiting. For example, any two or all of the training data creation device 10, the learning device 20, and the classification support device 30 may be configured as the same device. Furthermore, the trained model may be stored in an external device.
[0021] The functional configuration of the learning data creation device 10 will be described in more detail below. Fig. 3 is a block diagram showing the functional configuration of the learning data creation device 10. As shown in Fig. 3, the learning data creation device 10 includes a learning requirement specification text acquisition unit 101, a requirement specification classification table storage unit 102, a learning data generation unit 103, an operation input unit 104, a learning data storage unit 105, and a learning data output unit 106.
[0022] The training requirement specification text acquisition unit 101 acquires training requirement specification text (character strings) extracted from a training requirement specification document (semi-structured document) from an external device or storage medium, etc. The training requirement specification text acquisition unit 101 outputs the acquired training requirement specification text to the training data generation unit 103. Note that the training requirement specification text acquisition unit 101 may be configured to perform a process of extracting training requirement specification text from the training requirement specification document.
[0023] The requirement specification classification table storage unit 102 (classification storage unit) stores a pre-generated requirement specification classification table T1. The requirement specification classification table storage unit 102 is configured by a storage medium such as a RAM (Random Access Memory; a readable and writable memory), a flash memory, an EEPROM (Electrically Erasable Programmable Read Only Memory), or an HDD (Hard Disk Drive), or any combination of these storage media.
[0024] Here, an example of the table configuration of the required specification classification table T1 will be described. 4 is a diagram showing the table configuration of the requirements classification table T1. The requirements classification table T1 is a table showing a list of classification information that specifies the classification of the content of the requirements text (character strings) included in the requirements specification document (semi-structured document). As shown in FIG. 4, the requirements classification table T1 is tabular data in which at least four items, namely, "major classification," "minor classification," "classification type," and "threshold," are associated with each other.
[0025] "Major Category" is an item that indicates the major category of the learning requirements specification text. As shown in Figure 4, the values of "Major Category" include "Functional Requirements," "Non-Functional Requirements," and "Constraints." "Minor Category" is an item that indicates the minor category of the learning requirements specification text. As shown in Figure 4, when the value of "Major Category" is "Functional Requirements," no value is stored in "Minor Category." When the value of "Major Category" is "Non-Functional Requirements," the values of "Minor Category" include, for example, "Availability" and "Performance." When the value of "Major Category" is "Constraints," the values of "Minor Category" include, for example, "Development Environment" and "Operating Conditions."
[0026] "Classification type" is an item that indicates whether the classification identified by the "major classification" and "minor classification" is one of the predetermined classifications (i.e., the classification registered in the requirement specification classification table T1 when the classification support system 1 is in its initial state, for example), or whether it is a classification that has been added independently by, for example, a system administrator of the classification support system 1 after the system has started operating. As shown in Figure 4, the values of "classification type" are "basic" and "custom." When the value of "classification type" is "basic," this indicates that the classification identified by the "major classification" and "minor classification" is a predetermined classification. When the value of "classification type" is "custom," this indicates that the classification identified by the "major classification" and "minor classification" is a classification that has been added independently by a system administrator, for example.
[0027] The "threshold" is a value used to determine the classification result based on the probability output from the requirements classification model when classifying the requirements text by the classification support device 30. The "threshold" will be described in detail later.
[0028] The table structure of the requirement specification classification table T1 may be a structure in which the classification information is further hierarchically organized into multiple levels (for example, to include "medium classifications"), or it may be a non-hierarchical structure that has only "major classifications," for example.
[0029] Returning to FIG. 3, the explanation will be given again. The learning data generation unit 103 acquires the learning requirement specification text output from the learning requirement specification text acquisition unit 101. The learning data generation unit 103 also reads out the requirement specification classification table T1 stored in the requirement specification classification table storage unit 102. The learning data generation unit 103 then generates learning data D1 by associating the acquired learning requirement specification text with specific classification information selected by operation input from the operation input unit 104 from the classification information included in the read requirement specification classification table T1. Note that the classification information is information indicating a combination of a "major classification" value and a "minor classification" value.
[0030] Specifically, the learning data generation unit 103 displays, for example, the acquired learning requirement specification text and information indicating the read requirement classification specification table T1 on a display unit (not shown) such as a display provided in the learning data creation device 10. Then, an experienced engineer (user) with know-how on classifying requirement specification text performs an operation input via the operation input unit 104 to select and associate specific classification information corresponding to the input learning requirement specification text from the classification information included in the displayed requirement classification specification table T1. This generates learning data D1 in which the learning requirement specification text is associated with the specific classification information. The learning data generation unit 103 stores the generated learning data D1 in the learning data storage unit 105.
[0031] The operation input unit 104 receives operation inputs from a user (for example, an experienced engineer). The operation input unit 104 includes input devices such as a keyboard, a mouse, and a touch panel.
[0032] The learning data storage unit 105 stores the learning data D1 generated by the learning data generation unit 103. The learning data storage unit 105 stores, for example, several thousand to several hundred thousand pieces of learning data D1. The learning data storage unit 105 is configured by, for example, a storage medium such as a RAM, a flash memory, an EEPROM, or a HDD, or any combination of these storage media.
[0033] The learning data output unit 106 acquires the learning data D1 stored in the learning data storage unit 105. The learning data output unit 106 outputs the acquired learning data D1 to the learning device 20. For example, when generation of all the learning data D1 is completed, or when the number of pieces of generated learning data D1 reaches a predetermined threshold, the learning data output unit 106 outputs the learning data D1 to the learning device 20.
[0034] The functional configuration of the learning device 20 will be described in more detail below. Fig. 5 is a diagram showing the functional configuration of the learning device 20. As shown in Fig. 5, the learning device 20 includes a corpus acquisition unit 201, a preprocessing rule storage unit 202, a text preprocessing unit 203, a word vector learning unit 204, a trained model storage unit 205, a training data acquisition unit 206, a word vector conversion unit 207, a requirement specification classification learning unit 208, and a trained model output unit 209.
[0035] The corpus acquisition unit 201 acquires a corpus obtained from, for example, the Internet, etc. The corpus acquisition unit 201 outputs the acquired corpus to the text preprocessing unit 203. Note that instead of acquiring individual corpora, the corpus acquisition unit 201 may acquire a corpus group consisting of multiple corpora and perform processing to cut the corpus group into individual corpora.
[0036] The preprocessing rule storage unit 202 stores a preprocessing rule R1 in advance. The preprocessing rule R1 is information indicating rules used when preprocessing (e.g., natural language processing) is performed on character strings included in the corpus and on training requirement specification text included in the training data. For example, the preprocessing rule R1 includes rules for word segmentation, normalization, and deletion of unnecessary words. The preprocessing rule storage unit 202 is configured by, for example, a storage medium such as a RAM, a flash memory, an EEPROM, or a HDD, or any combination of these storage media.
[0037] The text pre-processing unit 203 acquires the corpus output from the corpus acquisition unit 201. The text pre-processing unit 203 also reads out the pre-processing rule R1 stored in the pre-processing rule storage unit 202. The text pre-processing unit 203 then performs pre-processing (e.g., natural language processing) on the character strings included in the acquired corpus based on the read out pre-processing rule R1. The text pre-processing unit 203 outputs the character strings included in the pre-processed corpus to the word vector training unit 204.
[0038] The word vector training unit 204 acquires character strings included in the preprocessed corpus output from the text preprocessing unit 203. The word vector training unit 204 performs machine learning using the acquired character strings as input to generate a word vector model M1, which is a training model of trained word vectors. The word vector training unit 204 stores the generated word vector model M1 in the trained model storage unit 205.
[0039] A word vector is a quantification (vectorization) of a word (feature) that is obtained by learning, using machine learning, the tendency of other words that are often used together with a word included in a sentence. Word2vec, for example, can be used as a method for converting into a word vector. The word vector model M1 is a learning model that has undergone machine learning to output a word vector in response to an input string of characters, for example.
[0040] The trained model storage unit 205 stores the word vector model M1 generated by the word vector learning unit 204 and the requirement specification classification model M2 (described later) generated by the requirement specification classification learning unit 208. The trained model storage unit 205 is configured by, for example, a storage medium such as a RAM, a flash memory, an EEPROM, or a HDD, or any combination of these storage media.
[0041] The learning data acquisition unit 206 acquires the learning data D1 output from the learning data creation device 10. The learning data acquisition unit 206 outputs the learning requirement text included in the acquired learning data D1 to the text preprocessing unit 203. In addition, the learning data acquisition unit 206 outputs the classification information included in the acquired learning data D1 to the requirement classification learning unit 208.
[0042] The text pre-processing unit 203 acquires the training requirement specification text output from the training data acquisition unit 206. The text pre-processing unit 203 also reads out the pre-processing rule R1 stored in the pre-processing rule storage unit 202. The text pre-processing unit 203 then performs pre-processing (e.g., natural language processing) on the acquired training requirement specification text based on the read out pre-processing rule R1. The text pre-processing unit 203 outputs the pre-processed training requirement specification text to the word vector conversion unit 207.
[0043] The word vector conversion unit 207 acquires the preprocessed training requirement specification text output from the text preprocessing unit 203. The word vector conversion unit 207 also reads out the word vector model M1 stored in the trained model storage unit 205. The word vector conversion unit 207 inputs the acquired training requirement specification text into the read-out word vector model M1, thereby converting the training requirement specification text into a word vector. The word vector conversion unit 207 outputs the training requirement specification text converted into a word vector to the requirement specification classification training unit 208.
[0044] The requirements classification learning unit 208 acquires classification information included in the learning data D1 output from the learning data acquisition unit 206. The requirements classification learning unit 208 also acquires the learning requirements text converted into word vectors output from the word vector conversion unit 207. The requirements classification learning unit 208 then performs machine learning using as training data the learning requirements text converted into the acquired word vectors and the acquired classification information, thereby obtaining a requirements classification model M2, which is a trained learning model.
[0045] The requirements specification classification model M2 is, for example, a neural network learning model that has been machine-learned to output information about classification results in response to input requirements specification text converted into word vectors. The requirements specification classification learning unit 208 stores the generated requirements specification classification model M2 in the learned model storage unit 205.
[0046] The trained model output unit 209 reads out the word vector model M1 and the requirement specification classification model M2 stored in the trained model storage unit 205. The trained model output unit 209 outputs the read out word vector model M1 and requirement specification classification model M2 to the classification assistance device 30.
[0047] The functional configuration of the classification support device 30 will be described in more detail below. Fig. 6 is a diagram showing the functional configuration of the classification support device 30. As shown in Fig. 6, the classification support device 30 includes a trained model acquisition unit 301, a trained model storage unit 302, a requirement specification text acquisition unit 303, a preprocessing rule storage unit 304, a text preprocessing unit 305, a word vector conversion unit 306, a requirement specification classification unit 307, a classification result analysis unit 308, and a classification result output unit 309.
[0048] The trained model acquisition unit 301 acquires the word vector model M1 and the requirement specification classification model M2 output from the learning device 20. The trained model acquisition unit 301 stores the word vector model M1 and the requirement specification classification model M2 in the trained model storage unit 302.
[0049] The trained model storage unit 302 stores the word vector model M1 and the requirement specification classification model M2 output from the trained model acquisition unit 301. The trained model storage unit 302 is configured by, for example, a storage medium such as RAM, flash memory, EEPROM, and HDD, or any combination of these storage media.
[0050] The requirements specification text acquisition unit 303 acquires requirements specification text (character strings) extracted from the requirements specification document (semi-structured document) to be classified from an external device or storage medium, etc. The requirements specification text acquisition unit 303 outputs the acquired requirements specification text to the text pre-processing unit 305. Note that the requirements specification text acquisition unit 303 may be configured to perform a process of extracting requirements specification text from the requirements specification document to be classified.
[0051] The preprocessing rule storage unit 304 stores the preprocessing rule R1 in advance. Note that the preprocessing rule storage unit 304 may acquire the preprocessing rule R1 from the preprocessing rule storage unit 202 of the learning device 20. The preprocessing rule storage unit 304 is configured by, for example, a storage medium such as RAM, flash memory, EEPROM, or HDD, or any combination of these storage media.
[0052] The text pre-processing unit 305 acquires the requirement specification text output from the requirement specification text acquisition unit 303. The text pre-processing unit 305 also reads out the pre-processing rule R1 stored in the pre-processing rule storage unit 304. The text pre-processing unit 305 then performs pre-processing (e.g., natural language processing) on the acquired requirement specification text based on the read out pre-processing rule R1. The text pre-processing unit 305 outputs the pre-processed requirement specification text to the word vector conversion unit 306.
[0053] The word vector conversion unit 306 acquires the preprocessed requirement specification text output from the text preprocessing unit 305. The word vector conversion unit 306 also reads out the word vector model M1 stored in the trained model storage unit 302. The word vector conversion unit 306 converts the requirement specification text into a word vector by inputting the acquired requirement specification text into the read-out word vector model M1. The word vector conversion unit 306 outputs the requirement specification text converted into the word vector to the requirement specification classification unit 307.
[0054] The requirement specification classification unit 307 acquires the requirement specification text converted into word vectors output from the word vector conversion unit 306. The requirement specification classification unit 307 also reads out the requirement specification classification model M2 stored in the trained model storage unit 302. The requirement specification classification unit 307 then inputs the acquired requirement specification text converted into word vectors into the read requirement specification classification model M2, thereby acquiring information on the classification result of the requirement specification text. The requirement specification classification unit 307 outputs information on the acquired classification result of the requirement specification text to the classification result analysis unit 308.
[0055] The information on the classification results of the requirement specification text output from the requirement specification classification unit 307 is, for example, a value indicating the probability of each of all classification information (i.e., combinations of "major classifications" and "minor classifications") included in the requirement specification classification table shown in Fig. 4. In other words, the information on the classification results of the requirement specification text is, for example, information indicating the probability that the requirement specification text to be classified belongs to each classification.
[0056] The classification result analysis unit 308 acquires information about the classification results of the requirement specification text output from the requirement specification classification unit 307. The classification result analysis unit 308 generates information indicating the classification results of the requirement specification text by analyzing the information about the classification results. The classification result analysis unit 308 outputs the generated information indicating the classification results of the requirement specification text to the classification result output unit 309.
[0057] The information indicating the classification result of the requirement specification text output from the classification result analysis unit 308 is, for example, information indicating a specific classification selected based on the value indicating the probability for each piece of classification information output from the requirement specification classification unit 307. For example, the information indicating the classification result is information indicating a classification whose probability value is equal to or greater than a predetermined threshold. Alternatively, for example, the information indicating the classification result is information indicating a classification corresponding to the top n (n is a natural number) values of the probability. Alternatively, for example, the information indicating the classification result is information indicating a classification corresponding to the largest value of the probability values.
[0058] The classification result analysis unit 308 may adjust (tune) the value of the "threshold" in the requirements specification classification table T1 shown in Fig. 4 based on the result of analyzing the information related to the classification results. For example, the classification result analysis unit 308 may update the value of the "threshold" so as to further improve the accuracy of classification of the requirements specification text. The adjustment (tuning) of the "threshold" value may be performed manually by an engineer (user) who has extensive experience in classifying requirements specification text, or may be performed automatically using a tuning tool or the like.
[0059] The classification result output unit 309 acquires information indicating the classification result output from the classification result analysis unit 308. The classification result output unit 309 outputs the acquired information indicating the classification result to an external device. Note that the classification result output unit 309 may be configured to display the information indicating the classification result on a display unit (not shown) such as a display provided in the classification support device 30.
[0060] An example of the operation of the learning data creation device 10 will now be described. 7 is a flowchart showing the operation of the learning data creating device 10. First, the learning data generating unit 103 reads out the required specification classification table T1 stored in the required specification classification table storage unit 102 (step S101).
[0061] The learning data generation unit 103 acquires the learning requirement specification text output from the learning requirement specification text acquisition unit 101 (step S102). The operation input unit 104 accepts operation input by the user (step S103). The learning data generation unit 103 then generates learning data D1 by associating the acquired learning requirement specification text with specific classification information selected by operation input from the operation input unit 104 from the classification information included in the read requirement specification classification table T1 (step S104). The learning data generation unit 103 stores the generated learning data D1 in the learning data storage unit 105 (step S105).
[0062] If the generation of learning data for all learning requirement specification texts has not been completed (i.e., if there is a learning requirement specification text to which classification information is not associated) (step S106, NO), the learning data creation device 10 continues to repeat the above-mentioned learning data generation process (steps S102 to S105). If the generation of learning data for all learning requirement specification texts has been completed (step S106, YES), the learning data output unit 106 outputs the learning data D1 stored in the learning data storage unit 105 to the learning device 20 (step S107). This completes the operation of the learning data creation device 10 shown in the flowchart of FIG. 7.
[0063] An example of the operation of the learning device 20 in generating the word vector model M1 will now be described. 8 is a flowchart showing the operation of the learning device 20 in generating the word vector model M1. First, the text preprocessing unit 203 reads out the preprocessing rule R1 stored in the preprocessing rule storage unit 202 (step S201).
[0064] The text pre-processing unit 203 acquires the corpus output from the corpus acquisition unit 201 (step S202). Then, the text pre-processing unit 203 performs pre-processing (e.g., natural language processing) on the character strings included in the acquired corpus based on the read pre-processing rule R1 (step S203).
[0065] The word vector training unit 204 acquires character strings included in the preprocessed corpus output from the text preprocessing unit 203 and stores them as training input data. If the training input data creation process has not been completed for all corpora (i.e., if there are corpora that have not been added to the training input data) (step S205: NO), the training device 20 continues to repeat the above-mentioned training input data creation process (steps S202 to S204).
[0066] When the process of creating training input data has been completed for all corpora (step S205: YES), the word vector training unit 204 performs machine learning using the training input data as input to generate a word vector model M1 and stores it in the trained model storage unit 205 (step S206). This completes the operation of the training device 20 in generating the word vector model M1, as shown in the flowchart in FIG. 8.
[0067] An example of the operation of the learning device 20 in generating the requirement specification classification model M2 will be described below. 9 is a flowchart showing the operation of the learning device 20 in generating the requirement specification classification model M2. First, the text preprocessing unit 203 reads out the preprocessing rule R1 stored in the preprocessing rule storage unit 202 (step S211).
[0068] The learning data acquisition unit 206 acquires the learning data D1 output from the learning data creation device 10 (step S212). The learning data acquisition unit 206 outputs the learning requirement text included in the acquired learning data D1 to the text preprocessing unit 203. In addition, the learning data acquisition unit 206 outputs the classification information included in the acquired learning data D1 to the requirement classification learning unit 208.
[0069] The text pre-processing unit 203 acquires the training requirement specification text output from the training data acquisition unit 206. Then, the text pre-processing unit 203 performs pre-processing (e.g., natural language processing) on the acquired training requirement specification text based on the read pre-processing rule R1 (step S213).
[0070] The word vector conversion unit 207 acquires the preprocessed training requirement specification text output from the text preprocessing unit 203. The word vector conversion unit 207 also reads out the word vector model M1 stored in the trained model storage unit 205. The word vector conversion unit 207 inputs the acquired training requirement specification text into the read out word vector model M1, thereby converting the training requirement specification text into a word vector (step S214).
[0071] The requirement specification classification learning unit 208 acquires classification information included in the learning data D1 output from the learning data acquisition unit 206. The requirement specification classification learning unit 208 also acquires the learning requirement specification text converted into word vectors output from the word vector conversion unit 207. Then, the requirement specification classification learning unit 208 stores data in which the learning requirement specification text converted into the acquired word vectors and the acquired classification information are associated with each other as learning input data (step S215).
[0072] If the process of creating learning input data has not been completed for all learning data (i.e., there is learning data that has not been added to the learning input data) (step S216: NO), the learning device 20 continues to repeat the above-mentioned process of creating learning input data (steps S212 to S215). If the process of creating learning input data has been completed for all learning data (step S216: YES), the requirements specification classification learning unit 208 performs machine learning using the learning input data as training data to generate a requirements specification classification model M2, which is a trained learning model, and stores it in the trained model storage unit 205 (step S217). The trained model output unit 209 outputs the word vector model M1 and the requirements specification classification model M2 stored in the trained model storage unit 205 to the classification support device 30 (step S218). This completes the operation of the learning device 20 in generating a requirements specification classification model, as shown in the flowchart of FIG. 9.
[0073] An example of the operation of the classification support device 30 will now be described. 10 is a flowchart showing the operation of the classification support device 30. First, the trained model acquisition unit 301 acquires the word vector model M1 and the requirement specification classification model M2 output from the learning device 20, and stores them in the trained model storage unit 302 (step S301). In addition, the text preprocessing unit 305 reads out the preprocessing rule R1 stored in the preprocessing rule storage unit 304 (step S302).
[0074] The text pre-processing unit 305 acquires the requirement specification text output from the requirement specification text acquisition unit 303 (step S303). Then, the text pre-processing unit 305 performs pre-processing (e.g., natural language processing) on the acquired requirement specification text based on the read pre-processing rule R1 (step S304).
[0075] The word vector conversion unit 306 acquires the preprocessed requirement specification text output from the text preprocessing unit 305. The word vector conversion unit 306 also reads out the word vector model M1 stored in the trained model storage unit 302. The word vector conversion unit 306 inputs the acquired requirement specification text into the read out word vector model M1, thereby converting the requirement specification text into a word vector (step S305).
[0076] The requirement specification classification unit 307 acquires the requirement specification text converted into word vectors output from the word vector conversion unit 306. The requirement specification classification unit 307 also reads out the requirement specification classification model M2 stored in the trained model storage unit 302. The requirement specification classification unit 307 then inputs the acquired requirement specification text converted into word vectors into the read requirement specification classification model M2, thereby executing requirement specification classification (step S306). As a result, the requirement specification classification unit 307 obtains information regarding the classification result of the requirement specification text.
[0077] The classification result analysis unit 308 acquires information related to the classification results of the requirement specification text output from the requirement specification classification unit 307. The classification result analysis unit 308 generates information indicating the classification results of the requirement specification text by analyzing the information related to the classification results (step S307). The classification result output unit 309 acquires information indicating the classification results output from the classification result analysis unit 308. The classification result output unit 309 outputs the acquired information indicating the classification results to an external device (step S308).
[0078] If requirement specification classification for all requirement specification texts has not been completed (i.e., if there is requirement specification text that has not been classified) (step S309: NO), the classification support device 30 continues to repeat the requirement specification classification process (steps S303 to S308). If requirement specification classification for all requirement specification texts has been completed (step S309: YES), the operation of the classification support device 30 shown in the flowchart of FIG. 10 ends.
[0079] According to at least one of the embodiments described above, a requirement specification classification learning unit 208 (or a storage unit) acquires learning data in which a requirement specification text for learning (learning information) based on character strings included in a requirement specification document for learning (semi-structured learning document) is associated with classification information that identifies the classification of the content of the requirement specification text for learning, inputs the requirement specification text for learning (learning input information) converted into a word vector based on the requirement specification text for learning (learning information), and the classification information, performs machine learning, and stores a requirement specification classification model M2 (trained model) indicating the trained learning model in a trained model storage unit 205 (storage unit). The system has a requirements specification classification unit 307 (classification unit) that acquires requirements specification text (classification information) based on character strings included in the requirements specification document (semi-structured document) to be classified, and acquires classification result-related information regarding the classification results of the contents of the requirements specification text (character strings) by inputting the requirements specification text (classification input information) converted into word vectors based on the requirements specification text (classification information) into the requirements specification classification model M2, and a classification result output unit 309 that outputs classification result information based on the classification result-related information, thereby making it possible to classify information included in semi-structured documents more efficiently without compromising the quality of the work.
[0080] Note that the training data creation device 10, the learning device 20, and the classification support device 30 in the above-described embodiments may be partially or entirely implemented by a computer. In this case, a program for implementing the control functions may be recorded on a computer-readable recording medium, and the program may be loaded and executed by a computer system. Note that the term "computer system" as used herein refers to the computer system built into the training data creation device 10, the learning device 20, and the classification support device 30, and includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computer systems. Furthermore, the term "computer-readable recording medium" may also include devices that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client. The program may be designed to implement some of the above-described functions, or may be capable of implementing the above-described functions in combination with a program already stored in the computer system.
[0081] Furthermore, the training data creation device 10, the learning device 20, and the classification support device 30 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the training data creation device 10, the learning device 20, and the classification support device 30 may be individually implemented as a processor, or some or all of them may be integrated into a processor. Furthermore, the integrated circuit implementation method is not limited to LSI, and may be implemented using a dedicated circuit or a general-purpose processor. Furthermore, if an integrated circuit implementation technology that can replace LSI emerges due to advances in semiconductor technology, an integrated circuit based on that technology may be used.
[0082] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents. [Explanation of symbols]
[0083] 1...Classification support system, 10...Learning data creation device, 20...Learning device, 30...Classification support device, 101...Learning requirement specification text acquisition unit, 102...Requirement specification classification table storage unit, 103...Learning data generation unit, 104...Operation input unit, 105...Learning data storage unit, 106...Learning data output unit, 201...Corpus acquisition unit, 202...Preprocessing rule storage unit, 203...Text preprocessing unit, 204...Word vector learning unit, 205... Trained model storage unit, 206...learning data acquisition unit, 207...word vector conversion unit, 208...requirements specification classification learning unit, 209...trained model output unit, 301...trained model acquisition unit, 302...trained model storage unit, 303...requirements specification text acquisition unit, 304...preprocessing rule storage unit, 305...text preprocessing unit, 306...word vector conversion unit, 307...requirements specification classification unit, 308...classification result analysis unit, 309...classification result output unit
Claims
1. a classification learning unit that acquires learning data in which learning information based on character strings contained in semi-structured learning documents that include at least text that can be classified into functional requirements that indicate requirements related to functions that must be installed in the system and non-functional requirements that indicate requirements other than the functional requirements, and classification information that identifies the classification of the content of the character strings, and performs machine learning by inputting learning input information based on the learning information and the classification information, and stores a learned model that indicates a learned learning model in a storage unit; a classification unit that acquires classification information based on character strings included in a semi-structured document that includes at least text that can be classified into the functional requirements and the non-functional requirements, and inputs classification input information based on the classification information into the trained model to classify the contents of the character strings into the functional requirements and the non-functional requirements, thereby acquiring classification result-related information that indicates a probability value for each classification; an analysis unit that generates classification result information indicating the classification result based on the classification result related information and a predetermined threshold; a classification result output unit that outputs the classification result information; Equipped with The analysis unit updates the threshold value based on a result of analyzing the classification result related information. Classification support system.
2. a classification storage unit that stores a classification table that lists classification information that identifies the classification of the content of character strings included in semi-structured documents; a training data generation unit that acquires character strings included in semi-structured training documents and generates training data in which the character strings are associated with specific categories selected from the category information included in the category table based on an operation input by a user; The classification assistance system of claim 1 further comprising:
3. a word vector conversion unit that inputs the training information into a word vector model that has been machine-learned using at least one of character strings included in the corpus and character strings included in the semi-structured training document as input, thereby acquiring the training input information that indicates the training information converted into word vectors. The classification support system according to claim 1 or 2, further comprising:
4. a word vector learning unit that performs machine learning using as input at least one of character strings included in the corpus and character strings included in the semi-structured learning document, and stores the word vector model representing a learned learning model in the storage unit; The classification assistance system of claim 3 further comprising:
5. The word vector conversion unit The classification information is input to the word vector model, thereby obtaining classification input information indicating the classification information converted into the word vector. The classification support system of claim 4.
6. The classification result information is information indicating a classification selected based on the probability value. The classification support system of claim 1 .
7. The classification result information is information indicating a plurality of classifications in which the probability value is equal to or greater than the threshold value. The classification support system of claim 6.
8. The trained model is a neural network learning model that has undergone machine learning to output the classification result related information in response to input of the classification input information. A classification support system according to any one of claims 1 to 7.
9. The word vector model is a learning model that has undergone machine learning to output the classification input information in response to the input of the classification information. The classification support system of claim 5.
10. The semi-structured learning document and the semi-structured document further include text that can be classified into constraints that indicate matters related to constraints during development and operation of the system. The classification support system of claim 1 .
11. a classification unit that acquires classification information based on character strings contained in a semi-structured document that includes at least text that can be classified into functional requirements that indicate requirements related to functions that must be installed in the system and non-functional requirements that indicate requirements other than the functional requirements, and inputs classification input information based on the classification information into a trained model to classify the contents of the character strings into the functional requirements and the non-functional requirements, thereby acquiring classification result-related information that indicates a probability value for each classification; an analysis unit that generates classification result information indicating the classification result based on the classification result related information and a predetermined threshold; a classification result output unit that outputs the classification result information; Equipped with The analysis unit updates the threshold value based on a result of analyzing the classification result related information. Classification support device.
12. a classification learning step in which a computer acquires learning data in which learning information based on character strings contained in semi-structured learning documents that include at least text that can be classified into functional requirements that indicate requirements related to functions that must be installed in the system and non-functional requirements that indicate requirements other than the functional requirements and classification information that identifies the classification of the content of the character strings is associated with each other, and the computer inputs learning input information based on the learning information and the classification information to perform machine learning, and stores a learned model that indicates the learned learning model in a memory unit; a classification step in which a computer acquires classification information based on character strings contained in a semi-structured document that includes at least text that can be classified into the functional requirements and the non-functional requirements, and inputs classification input information based on the classification information into the trained model to acquire classification result-related information that indicates a probability value for each classification, as information on the classification result of the classification of the contents of the character strings into the functional requirements and the non-functional requirements; an analyzing step of generating classification result information indicating the classification result based on the classification result related information and a predetermined threshold; a classification result output step of outputting the classification result information; an updating step of updating the value of the threshold based on a result of analyzing the classification result related information; A classification assistance method having the following.
13. On the computer, a classification learning step of acquiring learning data in which learning information based on character strings contained in semi-structured learning documents containing at least text that can be classified into functional requirements indicating requirements related to functions that must be installed in the system and non-functional requirements indicating requirements other than the functional requirements, and classification information that identifies the classification of the content of the character strings are associated with each other, inputting learning input information based on the learning information and the classification information to perform machine learning, and storing a learned model indicating the learned learning model in a storage unit; a classification step of acquiring classification information based on character strings contained in a semi-structured document including at least text that can be classified into the functional requirements and the non-functional requirements, and inputting classification input information based on the classification information into the trained model to acquire classification result-related information that indicates a probability value for each classification, as information on the classification result of the classification of the contents of the character strings into the functional requirements and the non-functional requirements; an analyzing step of generating classification result information indicating the classification result based on the classification result related information and a predetermined threshold; a classification result output step of outputting the classification result information; an updating step of updating the value of the threshold based on a result of analyzing the classification result related information; A program to execute.
Citation Information
Patent Citations
Required specification extraction method linked with architecture construction
JP2006127397A
Requirement definition support system based on natural language analysis, system designing support system, requirement definition support device, system designing support method, and program
JP2008250760A
Document classification system, document classification program, and document classification method
JP2011170786A
Requirement definition support system, method and program for data analysis
JP2012243194A
Project management support system
JP2014203228A