Source Code Pedigree Management via Contextual Grammar Parsing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current source code auditing tools face challenges in accurately identifying and managing copyright rights holders and their associated code snippets, particularly in complex development environments, leading to inaccurate results and excessive false positives due to the variability of copyright statements and the complexity of open source licensing.

Innovation Solution

A method and system for source code pedigree management that uses a contextual copyright grammar parser to identify copyright rights holders, filter out unauthorized constructs, and compile a list of valid and rejected copyright statements, focusing on source code that results in distributable binaries, with features like date and symbol recognition to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If pattern matching using regular expressions is used to identify copyright statements, then the search can be automated and applied to large codebases, but the accuracy decreases and false positives increase due to the variability of copyright statement formats

Engineering Contradiction:
Improveautomation of copyright statement identificationVSAvoidaccuracy of copyright statement identification
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary component - a contextual grammar parser - that mediates between the automated pattern matching process and the copyright statement identification. This parser acts as a bridge that transforms the automated search into accurate identification by validating matches against grammatical rules for copyright statements, thereby maintaining automation while improving precision

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of pattern matching from simple regular expressions to contextual grammar-based parsing. This parameter change allows the system to maintain automation while significantly improving accuracy by using grammatical context to validate whether a matched pattern is indeed a legitimate copyright statement, reducing false positives

Inventive Principle:
Principle #35Parameter changes

2Reliability

If every file in a development project is scanned for copyright statements, then comprehensive coverage is achieved, but the results become unwieldy and difficult to manage with excessive false positives

Engineering Contradiction:
Improvecomprehensive coverage of copyright identificationVSAvoidcomplexity of managing copyright identification results
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the relevant copyright information from the comprehensive file scan results. Instead of presenting all matches, the system extracts and displays only those copyright statements that are valid and relevant to the project, filtering out false positives and unnecessary information. This extraction approach maintains comprehensive coverage while simplifying the results for effective management

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the copyright identification process into distinct phases: initial comprehensive scanning, contextual validation using grammar parsers, and selective presentation of results. This segmentation allows the system to achieve thorough coverage while managing complexity by processing and filtering results in manageable stages rather than presenting all raw data at once

Inventive Principle:
Principle #1Segmentation

3Ease of manufacture

If simple pattern matching is used to identify copyright statements, then the implementation is straightforward and quick, but the results are inaccurate and cannot reliably identify rights holders

Engineering Contradiction:
Improveease of implementing copyright identificationVSAvoidaccuracy of rights holder identification
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces a contextual grammar parser as an intermediary layer between simple pattern matching and rights holder identification. This intermediary validates the pattern matches against grammatical rules for copyright statements and extracts the actual rights holder information, maintaining ease of implementation while dramatically improving accuracy in identifying the correct rights holders

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical system of simple regular expression matching with a more sophisticated contextual grammar-based parsing system. This substitution maintains the automated nature of the process (ease of manufacture) while improving precision by using grammatical context to accurately identify rights holders rather than relying solely on pattern matching

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8640101B2Pedigree analysis for software compliance management
Publication Date: 2014.01.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8640101B2 patent drawing
  • US8640101B2 patent drawing
  • US8640101B2 patent drawing

AI summary

Embodiments of the present invention address deficiencies of the art in respect to source code analysis and provide a novel and non-obvious method, system and computer program product for source code pedigree management. In one embodiment of the invention, a method for source code pedigree management can be provided. The method can include parsing source code to identify copyright rights holders for corresponding copyright constructs, rejecting copyright constructs not associated with corresponding rights holders, compiling a list of the identified copyright rights holders, corresponding copyright statements, and lists of files corresponding to each of the copyright rights holders, and displaying the compiled list.