Data Record Management Redundancy Detection via Frequency Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Database management systems face challenges in identifying and addressing non-meaningful content and redundancy in unstructured natural language data records, such as employee performance reviews, which can lead to inconsistent and inaccurate evaluations.
Innovation Solution
A system and method that processes unstructured natural language data records by generating a structured dataset, normalizing frequency values using inverse proportionality, determining redundancy prediction values through cosine similarity, and identifying records with excessive similarity or focus on a single topic, thereby flagging potentially copied or generic content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If unstructured natural language data records are stored and analyzed, then comprehensive evaluation data is available, but redundancy and non-meaningful content cannot be effectively identified
Solution Approach 1:
The patent transforms the unstructured natural language data into a structured dataset with frequency values, then applies parameter transformation through inverse proportionality normalization. This converts raw frequency counts into normalized values that enable meaningful comparison and redundancy detection across different data records, directly addressing the inability to identify redundancy in unstructured data.
Solution Approach 2:
The patent replaces manual or simple mechanical text comparison with a computational system that uses cosine similarity calculations and inverse proportionality transformations. This substitution enables automated, precise identification of redundant content by computing mathematical similarities between normalized frequency vectors, significantly improving evaluation accuracy.
2Quantity of substance
If frequency values are used to represent term importance, then term occurrence is captured, but redundancy detection is insufficient due to lack of normalization
Solution Approach 1:
The patent applies inverse proportionality transformation to convert raw frequency values into normalized values. This parameter change ensures that terms appearing in all records (high frequency) receive low normalized values, while terms appearing in fewer records receive higher normalized values. This transformation enables reliable redundancy detection by highlighting terms that are uniquely important to specific records rather than universally common.
3Ease of manufacture
If unstructured data is processed without structuring, then data collection is simple, but analysis precision and redundancy identification are poor
Solution Approach 1:
The patent segments the unstructured natural language data into individual terms and represents each term's occurrence through frequency values in a structured dataset. This segmentation transforms continuous unstructured text into discrete, analyzable units, enabling precise measurement and comparison while maintaining the ease of initial data collection from various sources.
Data Source
AI summary
A system and method of data record management is provided. The system comprises a processor and a memory coupled to the processor that stores processor-executable instructions that when executed configure the processor to perform the method. The method comprises receiving a plurality of unstructured natural language data records, generating a structured dataset based on the plurality of unstructured natural language data records, transforming the structured dataset to normalize the respective frequency values based on inverse proportionality of the respective frequency values, determining a redundancy prediction value associated with that unstructured natural language data record based on the transformed structured dataset, and displaying on a graphical user interface a message identifying one or more unstructured natural language data records being associated with a redundancy prediction value greater than a threshold value. The structured dataset includes a frequency value associated with respective terms of each of the plurality of unstructured natural language data records.


