Data Repository Platform for File Similarity Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches make it difficult to find and reuse supporting materials related to specific products or components, leading to wasted time in recreating these materials.
Innovation Solution
A data repository management platform that receives input strings with identifiers, searches for corresponding files, computes similarities, ranks files, groups them based on identifiers, and generates logical divisions for easy access.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If supporting materials are stored in data repositories using current approaches, then files are stored in logical drives, but files become difficult to find and compile, leading to loss of time
Solution Approach 1:
The system pre-computes similarity scores between files and potential search queries, and pre-organizes files into logical divisions based on their content characteristics. This preliminary organization allows users to quickly retrieve relevant supporting materials without manually searching through entire logical drives, directly reducing the time loss and improving operational ease.
2Productivity
If supporting materials are not re-used, then new materials must be created each time, but this leads to wasted time and redundant work
Solution Approach 1:
The system enables users to copy existing supporting materials from the data repository by presenting relevant files based on similarity matching. Users can retrieve previously created materials and reuse them for new projects, eliminating redundant creation work and significantly improving productivity while reducing time loss.
3Ease of operation
If files are organized in traditional logical drives, then storage is simple, but searching and retrieving relevant files becomes difficult
Solution Approach 1:
The system adds a semantic dimension to traditional logical drive organization by computing similarity scores based on file content, metadata, and relationships. Files are organized not only by their physical location in logical drives but also by their semantic similarity to search queries, creating a multi-dimensional organization system that improves retrieval ease without adding significant complexity.
Data Source
AI summary
A method comprises receiving an input string comprising one or more identifiers for an operation, and searching at least one data repository for files corresponding to the one or more identifiers. Similarities between the input string and respective ones of the files corresponding to the one or more identifiers are computed, and the files corresponding to the one or more identifiers are ranked based on the computed similarities. The method further comprises grouping at least a portion of the ranked files into at least one group based on the one or more identifiers. At least one division in a logical drive of the at least one data repository is generated, wherein the at least one division corresponds to the at least one group and comprises at least the portion of the ranked files of the at least one group.


