Software Identification from Unstructured Text via Token Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods struggle to accurately identify unique software from unstructured text, leading to difficulties in managing software installations across networked devices due to irregularities and ambiguities in text formats, which complicates tasks like vulnerability assessment and license management.

Innovation Solution

A computing device processes unstructured text by dividing it into tokens, classifying them using a classification dataset, matching permutations to unique software identifiers defined by a common structured format like CPE, and selecting a single identifier according to set rules, generating text in a standardized format to indicate installed software.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If unstructured text is processed using existing methods, then software identification is attempted, but accuracy is poor due to irregularities and ambiguities in text formats

Engineering Contradiction:
Improvesoftware identification accuracyVSAvoidtext processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments unstructured text into discrete tokens and classifies each token into specific software parameters (product name, version, vendor, etc.). This segmentation transforms ambiguous text into structured data that can be accurately matched to software identifiers, resolving the contradiction between handling complex unstructured text and achieving accurate identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary classification dataset that maps text tokens to software parameters. This intermediary layer acts as a bridge between unstructured text and structured software identifiers, enabling accurate software identification without requiring complex direct parsing of ambiguous text formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual methods are used to identify software from unstructured text, then flexibility is maintained, but productivity is low and time-consuming

Engineering Contradiction:
Improvesoftware identification speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary classification of text tokens into software parameters before matching to software identifiers. This preliminary action organizes the data in advance, enabling rapid automated matching and significantly improving productivity while reducing the time loss associated with manual software identification processes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service automated software identification by processing unstructured text through classification and matching algorithms without human intervention. This automation eliminates manual processing bottlenecks, dramatically increasing productivity and reducing time loss while maintaining accuracy through systematic rule-based matching.

Inventive Principle:
Principle #25Self-service

3Reliability

If multiple software identifiers are generated from unstructured text, then comprehensive matching is achieved, but difficulty in selecting the correct identifier increases

Engineering Contradiction:
Improvesoftware identification reliabilityVSAvoididentifier selection ease
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements feedback mechanisms through a set of selection rules that evaluate multiple matched software identifiers and select the most appropriate one. The system uses classification confidence scores and parameter matching quality as feedback to automatically resolve ambiguities, improving reliability while maintaining ease of operation by eliminating manual selection requirements.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes parameters by transforming unstructured text into classified software parameters with specific attributes (product, version, vendor, etc.). This parameter transformation enables systematic comparison and selection among multiple identifiers using defined criteria, making the selection process both reliable and operationally simple through automated rule-based decision making.

Inventive Principle:
Principle #35Parameter changes

4Stability of the object's composition

If standardized format output is generated, then consistency is improved, but processing complexity increases due to permutation matching

Engineering Contradiction:
Improveoutput format consistencyVSAvoidmatching process complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent segments the matching process into discrete steps: token classification into parameters, parameter permutation generation, and systematic matching against standardized identifiers. This segmentation manages processing complexity by breaking down the complex permutation matching into manageable, automated steps while ensuring consistent standardized output format through structured data transformation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11995424B2Identifying unique computer software in unstructured text
Publication Date: 2024.05.28 AXONIUS SOLUTIONS LTD
  • US11995424B2 patent drawing
  • US11995424B2 patent drawing
  • US11995424B2 patent drawing

AI summary

There is provided a computing device for identifying unique software installed on network connected devices, comprising: a processor executing a code for: for each unstructured text for the network connected devices, wherein the unstructured texts are extracted by different code sensors from different applications, wherein each unstructured text indicates an identity of software installed on device(s): dividing the unstructured text into token(s), classifying tokens to software parameter(s) using classification dataset(s), matching subsets of permutations of the software parameters and corresponding tokens to unique software identifiers defined by a common structured format, selecting one unique software identifier according to a set of rules, and generating a text satisfying the common structured format, the text indicating unique software installed on each device.