Software Identification from Unstructured Text via Token Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately identify unique software from unstructured text, leading to difficulties in managing software installations across networked devices due to irregularities and ambiguities in text formats, which complicates tasks like vulnerability assessment and license management.
Innovation Solution
A computing device processes unstructured text by dividing it into tokens, classifying them using a classification dataset, matching permutations to unique software identifiers defined by a common structured format like CPE, and selecting a single identifier according to set rules, generating text in a standardized format to indicate installed software.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If unstructured text is processed using existing methods, then software identification is attempted, but accuracy is poor due to irregularities and ambiguities in text formats
Solution Approach 1:
The patent segments unstructured text into discrete tokens and classifies each token into specific software parameters (product name, version, vendor, etc.). This segmentation transforms ambiguous text into structured data that can be accurately matched to software identifiers, resolving the contradiction between handling complex unstructured text and achieving accurate identification.
Solution Approach 2:
The patent introduces an intermediary classification dataset that maps text tokens to software parameters. This intermediary layer acts as a bridge between unstructured text and structured software identifiers, enabling accurate software identification without requiring complex direct parsing of ambiguous text formats.
2Productivity
If manual methods are used to identify software from unstructured text, then flexibility is maintained, but productivity is low and time-consuming
Solution Approach 1:
The patent performs preliminary classification of text tokens into software parameters before matching to software identifiers. This preliminary action organizes the data in advance, enabling rapid automated matching and significantly improving productivity while reducing the time loss associated with manual software identification processes.
Solution Approach 2:
The system enables self-service automated software identification by processing unstructured text through classification and matching algorithms without human intervention. This automation eliminates manual processing bottlenecks, dramatically increasing productivity and reducing time loss while maintaining accuracy through systematic rule-based matching.
3Reliability
If multiple software identifiers are generated from unstructured text, then comprehensive matching is achieved, but difficulty in selecting the correct identifier increases
Solution Approach 1:
The patent implements feedback mechanisms through a set of selection rules that evaluate multiple matched software identifiers and select the most appropriate one. The system uses classification confidence scores and parameter matching quality as feedback to automatically resolve ambiguities, improving reliability while maintaining ease of operation by eliminating manual selection requirements.
Solution Approach 2:
The patent changes parameters by transforming unstructured text into classified software parameters with specific attributes (product, version, vendor, etc.). This parameter transformation enables systematic comparison and selection among multiple identifiers using defined criteria, making the selection process both reliable and operationally simple through automated rule-based decision making.
4Stability of the object's composition
If standardized format output is generated, then consistency is improved, but processing complexity increases due to permutation matching
Solution Approach 1:
The patent segments the matching process into discrete steps: token classification into parameters, parameter permutation generation, and systematic matching against standardized identifiers. This segmentation manages processing complexity by breaking down the complex permutation matching into manageable, automated steps while ensuring consistent standardized output format through structured data transformation.
Data Source
AI summary
There is provided a computing device for identifying unique software installed on network connected devices, comprising: a processor executing a code for: for each unstructured text for the network connected devices, wherein the unstructured texts are extracted by different code sensors from different applications, wherein each unstructured text indicates an identity of software installed on device(s): dividing the unstructured text into token(s), classifying tokens to software parameter(s) using classification dataset(s), matching subsets of permutations of the software parameters and corresponding tokens to unique software identifiers defined by a common structured format, selecting one unique software identifier according to a set of rules, and generating a text satisfying the common structured format, the text indicating unique software installed on each device.


