AI Content Citation Retrieval for Copyright Similarity Screening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generative AI tools often fail to provide information about the base content items used for generating output content, leading to potential copyright infringement and uncertainty regarding the originality and suitability of the output content.
Innovation Solution
A citator system searches content databases using text prompt embeddings to generate citations for output content, identifying similar base content items and providing metadata such as author, copyright information, and similarity scores, with the option to alter the output content to reduce similarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If generative AI tools create output content using trained ML models that draw from base content items, then the output content can be generated creatively, but the system fails to provide information about which base content items were used, leading to copyright infringement risks
Solution Approach 1:
The system performs preliminary actions by searching for base content items and generating citations before the output content is finalized. The citator system proactively identifies potential copyright issues and provides remediation options in advance, allowing users to address infringement risks before deployment.
Solution Approach 2:
The patent introduces a citator system as an intermediary component between the generative AI model and the output content. This mediator searches content databases, identifies similar base content items, generates citations, and provides remediation guidance, thereby bridging the information gap without requiring fundamental changes to the core ML model.
2Measurement precision
If the output content is generated to be creative and original, then the technological sophistication is improved, but the user cannot know the extent of originality or identify similar base content items
Solution Approach 1:
The system implements feedback mechanisms by providing users with citation information that indicates the degree of similarity between output content and base content items. The citator system returns feedback in the form of citations that include similarity metrics, allowing users to assess originality and make informed decisions about usage.
3Measurement precision
If the user knows the base content used for output content and estimates the overlap, then copyright risk awareness is improved, but the user cannot automatically identify specific portions of output content that are similar to base content
Solution Approach 1:
The patent implements self-service capabilities where the citator system automatically performs the complex tasks of searching content databases, comparing output content with base content items, generating citations, and identifying similar portions. This eliminates the need for users to manually perform these analysis tasks.
4Reliability
If the system provides detailed citation information about base content items, then copyright compliance is improved, but the processing time and computational resources increase
Solution Approach 1:
The system applies partial action by searching content databases and generating citations only when necessary, rather than continuously analyzing all possible base content items. The citator system performs selective searching and comparison to provide sufficient copyright compliance information without excessive computational resources.
Data Source
AI summary
A citation for output content that is generated by a trained generative machine learning (ML) model is disclosed. A content database is filtered based on a text prompt embedding generated based on the same text prompt input to the ML model to generate the output content. Further filtering may be performed using an output content embedding generated based on the output content generated by the ML model. A base content item is then estimated as being similar to the output content generated by the ML model by filtering the content list using component/content features generated based on the output content. A similarity score is generated and the citation identifying the base content item is provided to the ML model. In response to determining that the similarity meets a first threshold similarity criterion, an alternative output content may be generated with or without further user input.


