Content File Similarity Detection for Copying Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In collaborative software development environments, detecting copied software elements is challenging due to the large number of constituent software elements and multiple versions, leading to an overwhelming number of potential sources, which frustrates users and increases the risk of legal rights infringement.

Innovation Solution

A method and system that automatically identify similar content in content files by examining portions of the files, determining supersets, and providing an indication of the best match based on additional information such as licensing, date, and directory structure, to narrow down the sources of potential copying.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a database of constituent software elements is used to identify copied sections, then the ability to detect copying is improved, but the number of matches becomes unwieldy large when multiple versions and copies exist

Engineering Contradiction:
Improvecopying detection capabilityVSAvoidnumber of matches
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent replaces manual inspection of multiple matches with an automated system that uses fingerprinting and similarity algorithms to objectively determine the best match. The system automatically compares constituent elements against the aggregated product, generates fingerprints, and ranks matches based on similarity metrics, eliminating the need for users to manually evaluate numerous potential sources.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple versions and copies of constituent software elements are included in the database, then comprehensive copying detection is achieved, but user frustration increases due to inability to efficiently narrow down sources

Engineering Contradiction:
Improvecomprehensive copying detectionVSAvoiduser ability to narrow down sources
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system provides feedback by ranking matches based on similarity metrics and presenting the best match first. The automated comparison and ranking process gives users immediate, actionable information about which constituent element is most likely the source of copying, allowing them to efficiently verify and take appropriate measures without being overwhelmed by multiple unranked matches.

Inventive Principle:
Principle #23Feedback

3Reliability

If a large number of constituent software elements are compared against the aggregated product, then thorough copying detection is achieved, but the complexity of the comparison process increases

Engineering Contradiction:
Improvethorough copying detectionVSAvoidcomparison process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the comparison process into distinct steps: generating fingerprints from constituent elements, comparing fingerprints against the aggregated product, ranking matches based on similarity, and presenting results. This segmentation simplifies the overall complexity by breaking down the large-scale comparison into manageable, automated stages that can be processed systematically.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8010538B2Methods and systems for reporting regions of interest in content files
Publication Date: 2011.08.30 BLACK DUCK SOFTWARE INC
  • US8010538B2 patent drawing
  • US8010538B2 patent drawing
  • US8010538B2 patent drawing

AI summary

A content file is examined and compared against one or more comparison files. An indication is provided that the content file is similar to the one comparison file that is the best match with the examined content file.