Automated Requirement Clustering via Semantic Similarity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual grouping of system requirements in text documents is labor-intensive and dependent on user knowledge, making it inefficient for large or complex systems.
Innovation Solution
A device and method that utilize natural language processing to identify themes and requirements, performing similarity analysis to determine dominant themes and terms, which are then clustered to facilitate system design and development.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual grouping of system requirements is performed, then user knowledge and control are utilized, but labor intensity and time consumption increase significantly
Solution Approach 1:
The patent replaces manual mechanical grouping operations with automated natural language processing and similarity analysis algorithms. The system automatically extracts requirements from text documents, computes similarity scores between requirements and themes, and generates groupings without human intervention, thereby eliminating labor-intensive manual work while maintaining grouping quality through algorithmic precision.
Solution Approach 2:
The system enables self-service automated requirement grouping by utilizing the inherent semantic information within the requirements themselves. The similarity analysis algorithm autonomously determines groupings based on textual content without requiring external expert knowledge or manual input, allowing the system to serve itself in performing the grouping task.
2Reliability
If manual grouping of system requirements is performed, then user knowledge is applied, but efficiency decreases for large or complex systems
Solution Approach 1:
The patent substitutes manual expert-based grouping with automated natural language processing systems that can handle large volumes of requirements efficiently. The similarity analysis algorithm processes numerous requirements simultaneously, computing similarity scores and generating groupings at speeds impossible for manual operation, thereby dramatically improving productivity while maintaining reliability through consistent algorithmic application.
Solution Approach 2:
The system achieves universality by creating a general-purpose automated grouping mechanism that can handle various types of system requirements across different domains. The natural language processing and similarity analysis approach is domain-agnostic and can process requirements from diverse systems without requiring domain-specific manual intervention, thereby improving both efficiency and reliability across multiple applications.
3Productivity
If automated natural language processing is used to group requirements, then manual effort is reduced, but dependency on algorithmic accuracy increases
Solution Approach 1:
The patent replaces simple manual grouping operations with sophisticated automated natural language processing systems. While this increases algorithmic complexity, it eliminates the need for manual expert knowledge and enables high-speed automated processing. The complexity is justified by the significant productivity gains and the ability to handle large-scale requirement datasets that would be intractable manually.
Data Source
AI summary
A device may analyze text to identify a set of text portions of interest, and may analyze the text to identify a set of terms included in the set of text portions. The device may perform a similarity analysis to determine a similarity score. The similarity score may be determined between each term, included in the set of terms, and each text portion, included in the set of text portions, or the similarity score may be determined between each term and each other term included in the set of terms. The device may determine a set of dominant terms based on performing the similarity analysis. The set of dominant terms may include at least one term with a higher average degree of similarity than at least one other term. The device may provide information that identifies the set of dominant terms.


